arXiv Papers

topic · knowledge/papers-arxiv
DOC.
knowledge/papers-arxiv
REV.
596 evt
DATE.
29-MAY-2026
SCOPE.
custom
§01

about

Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 507 events in this window (596 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01Scaling Native Multimodal Pre-Training From ScratchAlthough large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby ach{"url":"https://arxiv.org/abs/2607.22043","tldr":"Although large language models…
EVENT. cms2w1swID. cms2w1swa9pjbkh0c3szxpqhxSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.22043",
  "tldr": "Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain ",
  "title": "Scaling Native Multimodal Pre-Training From Scratch",
  "authors": [
    "Haoyuan Wu",
    "Aoqi Wu",
    "Hai Wang",
    "Jiajia Wu",
    "Jinxiang Ou",
    "Bei Yu"
  ],
  "upvotes": 7,
  "arxiv_id": "2607.22043",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Tencent Hunyuan",
  "project_page": null,
  "published_at": "2026-07-24T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}
02Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative RenderingRecent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that cont{"url":"https://arxiv.org/abs/2607.21848","tldr":"Recent conditional video gener…
EVENT. cms2w1s5ID. cms2w1s589pj9kh0cp716z9nuSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21848",
  "tldr": "Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after",
  "title": "Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering",
  "authors": [
    "Wenchao Ma",
    "Changran Liu",
    "Sharon X. Huang",
    "Haomiao Jiang"
  ],
  "upvotes": 2,
  "arxiv_id": "2607.21848",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://wenchao-m.github.io/ClosetheLoop.github.io/#",
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}
03LAMAR: An Open Language-Aware Multilingual Alignment RerankerIn multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document language when ordering semantical{"url":"https://arxiv.org/abs/2607.22042","tldr":"In multilingual retrieval augm…
EVENT. cms2w1rjID. cms2w1rjn9pj7kh0c3a37mxcoSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.22042",
  "tldr": "In multilingual retrieval augmented generation, a retriever can retrieve relevant documents written in multiple languages, which are subsequently reranked before answer generation. However, it remains unclear whether existing multilingual rerankers consider document language when ordering semantically relevant candidates. Our analysis shows that these rerankers do not consistently prioritize documents written in the same language as the query when semantically equivalent documents are available ",
  "title": "LAMAR: An Open Language-Aware Multilingual Alignment Reranker",
  "authors": [
    "Seongtae Hong",
    "Youngjoon Jang",
    "Jungseob Lee",
    "Seungyoon Lee",
    "Heuiseok Lim"
  ],
  "upvotes": 2,
  "arxiv_id": "2607.22042",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "NLP & AI - Korea University",
  "project_page": "https://huggingface.co/nlpai-lab/LAMAR-600m",
  "published_at": "2026-07-24T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}
04IDEAgent: Agentic Quality-Diversity Search for Research Idea GenerationLarge Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in c{"url":"https://arxiv.org/abs/2607.22375","tldr":"Large Language Models (LLMs) h…
EVENT. cms2w1qyID. cms2w1qy89pj5kh0czbzdewdySRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.22375",
  "tldr": "Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation should be treated as a conjunction of both objectives an",
  "title": "IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation",
  "authors": [
    "Varun Gumma",
    "Navonil Majumder",
    "Soumitra Sinhahajari",
    "Soujanya Poria"
  ],
  "upvotes": 3,
  "arxiv_id": "2607.22375",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Deep Cognition and Language Research (DeCLaRe) Lab",
  "project_page": null,
  "published_at": "2026-07-24T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}
05Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture ProblemsProduction AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history wh{"url":"https://arxiv.org/abs/2607.21503","tldr":"Production AI agents' failures…
EVENT. cms2w1qcID. cms2w1qcl9pj3kh0czawxcekpSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21503",
  "tldr": "Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, large tool definitions, and ballooning tool outputs. Agents drown in their own accumulating history while paying a token cost that grows every turn, producing missing recalls within and across conversations. The incumbent response treats this as a storage-and-retrieval problem. We argue that framing i",
  "title": "Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems",
  "authors": [
    "Gaurav Dadhich"
  ],
  "upvotes": 5,
  "arxiv_id": "2607.21503",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/6a6384d16161ae6033bb5f69/H2Ez-V-zKCokNWACm_I0x.png",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://maximem.ai/synap",
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}
06OpenForgeRL: Train Harness-native Agents in Any EnvironmentModern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks can{"url":"https://arxiv.org/abs/2607.21557","tldr":"Modern AI agents rely on elabo…
EVENT. cmrzoaulID. cmrzoaul38xzrkh0c3hlcanv0SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21557",
  "tldr": "Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex harnesses also make agents hard to train end-to-end with open infrastructure, whose SFT/RL stacks cannot natively express stateful, multi-process harness inference. To address this, we present OpenForgeRL, an open-source framework for training harness-based agents end-to-end in diverse environments. ",
  "title": "OpenForgeRL: Train Harness-native Agents in Any Environment",
  "authors": [
    "Xiao Yu",
    "Baolin Peng",
    "Ruize Xu",
    "Hao Zou",
    "Qianhui Wu",
    "Hao Cheng",
    "Wenlin Yao",
    "Nikhil Singh",
    "Zhou Yu",
    "Jianfeng Gao"
  ],
  "upvotes": 3,
  "arxiv_id": "2607.21557",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Microsoft",
  "project_page": null,
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}
07Self-Supervised Learning of Structured Dynamics from VideosUnderstanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in nat{"url":"https://arxiv.org/abs/2607.21576","tldr":"Understanding motion in video …
EVENT. cmryykk8ID. cmryykk8z8qpjkh0c3ikm81uuSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21576",
  "tldr": "Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We ",
  "title": "Self-Supervised Learning of Structured Dynamics from Videos",
  "authors": [
    "Lukas Knobel",
    "Andrew Zisserman",
    "Yuki M. Asano"
  ],
  "upvotes": 12,
  "arxiv_id": "2607.21576",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Fundamental AI Lab at UTN",
  "project_page": "https://lukasknobel.github.io/projects/StructuredDynamics",
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}
08SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video GenerationWe introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence {"url":"https://arxiv.org/abs/2607.21553","tldr":"We introduce SANA-Video 2.0, a…
EVENT. cmryykjjID. cmryykjjn8qphkh0c8bohniweSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21553",
  "tldr": "We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favorable long-sequence scaling of linear attention. To avoid quadratic attention throughout, Hybrid Linear-Softmax Attention combines gated linear attention for O(N)-dominated mixing with periodic gated-softmax anchors at a",
  "title": "SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation",
  "authors": [
    "Junsong Chen",
    "Jincheng Yu",
    "Yitong Li",
    "Shuchen Xue",
    "Haozhe Liu",
    "Jingyu Xin",
    "Yuyang Zhao",
    "Tian Ye",
    "Zhangjie Wu",
    "Zian Wang",
    "Daquan Zhou",
    "Ping Luo"
  ],
  "upvotes": 9,
  "arxiv_id": "2607.21553",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/645b5b09bc7518912e1f9733/8WUF7Xm9QNEHHKgPGrfNO.mp4",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "NVIDIA",
  "project_page": "https://nvlabs.github.io/Sana/Video2",
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}
09ICAE-Bench: Evaluating Coding Agents as Interactive Project BuildersThe recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities includin{"url":"https://arxiv.org/abs/2607.21217","tldr":"The recent emergence of vibe-c…
EVENT. cmrylpn6ID. cmrylpn6w8ncfkh0cw49lgcdkSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21217",
  "tldr": "The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirement clarification, tool use, debugging, and repository-level construction. Yet existing benchmarks have not fully caught up with this shift, evaluating agents on static, fully spec",
  "title": "ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders",
  "authors": [
    "Zhongyuan Peng",
    "Dan Huang",
    "Chuyu Zhang",
    "Caijun Xu",
    "Changyi Xiao",
    "Shibo Hong",
    "David Lo",
    "Lin Qiu",
    "Xuezhi Cao",
    "Jiyuan He",
    "Yixin Cao"
  ],
  "upvotes": 4,
  "arxiv_id": "2607.21217",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
10Streaming Multi-Agent Autoregressive Diffusion Model with World State RegistersMulti-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared st{"url":"https://arxiv.org/abs/2607.21594","tldr":"Multi-agent interactive world …
EVENT. cmrylpmhID. cmrylpmhg8ncdkh0cutpgy76nSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21594",
  "tldr": "Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registe",
  "title": "Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers",
  "authors": [
    "Sicheng Mo",
    "Yuheng Li",
    "Ziyang Leng",
    "Krishna Kumar Singh",
    "Bolei Zhou"
  ],
  "upvotes": 3,
  "arxiv_id": "2607.21594",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/XVXjWFpcUXMn16V5Bstnl.qt",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://vail-ucla.github.io/worldweaver/",
  "published_at": "2026-07-23T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}
showing 1–10 of 507older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above