arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 498 events in this window (598 total on topic). adjust the range or clear it with ALL.
range
01Streaming Multi-Agent Autoregressive Diffusion Model with World State RegistersMulti-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared st{"url":"https://arxiv.org/abs/2607.21594","tldr":"Multi-agent interactive world …
EVENT. cmrylpmhID. cmrylpmhg8ncdkh0cutpgy76nSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21594",
"tldr": "Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video diffusion pipelines carry forward observation history as conditioning context, which makes shared state difficult to maintain in multi-agent and multi-view settings. We present WorldWeaver (W^2), a streaming multi-agent video diffusion model that augments rollout with cross-agent world state registe",
"title": "Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers",
"authors": [
"Sicheng Mo",
"Yuheng Li",
"Ziyang Leng",
"Krishna Kumar Singh",
"Bolei Zhou"
],
"upvotes": 3,
"arxiv_id": "2607.21594",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/XVXjWFpcUXMn16V5Bstnl.qt",
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://vail-ucla.github.io/worldweaver/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}02TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable ManipulationThe development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently{"url":"https://arxiv.org/abs/2607.21017","tldr":"The development of generalizab…
EVENT. cmrylplrID. cmrylplrw8ncbkh0cuogawg9fSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21017",
"tldr": "The development of generalizable robotic manipulation policies is inherently bounded by the availability of large-scale, high-fidelity scene data. While recent automated synthesis methods attempt to bridge this gap via text-to-layout hallucination or simplified procedural generation, they frequently suffer from physical implausibility and fail to capture the complex, dense clutter of actual human environments. In this paper, we introduce TableVerse, a fully automated Real2Sim pipeline that shift",
"title": "TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation",
"authors": [
"Boyuan Wang",
"Yue Zhang",
"Xutao Xue",
"Xueyu Song",
"Yu Sun"
],
"upvotes": 3,
"arxiv_id": "2607.21017",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "ByteDance",
"project_page": "https://bytedance.github.io/TableVerse/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}03NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsTraditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its method{"url":"https://arxiv.org/abs/2607.20709","tldr":"Traditional agent development …
EVENT. cmrylpl2ID. cmrylpl2f8nc9kh0czoswq8g6SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20709",
"tldr": "Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of \"...\" is completed at runtime by an",
"title": "NVIDIA-labs OO Agents: Native Python Object-Oriented Agents",
"authors": [
"Paul Furgale",
"Severin Klingler",
"James Nolan",
"Matt Staats",
"Gaia Di Lorenzo",
"Elisa Martinez Abad",
"Christian Schüller",
"Razvan Dinu",
"Alessio Devoto",
"Pascal Berard",
"Gal Kaplun",
"Elad Sarafian"
],
"upvotes": 4,
"arxiv_id": "2607.20709",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "NVIDIA",
"project_page": null,
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}04Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task ConstructionWe introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent {"url":"https://arxiv.org/abs/2607.20911","tldr":"We introduce Tencent WorkBuddy…
EVENT. cmrylpkcID. cmrylpkcp8nc7kh0cunqvc8jpSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20911",
"tldr": "We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and re",
"title": "Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction",
"authors": [
"Tencent WorkBuddy Bench Team",
"Siqi Cai",
"Shaopeng Chen",
"Xiang Fei",
"Yong Mao",
"Zihan Xu",
"Zhiheng Lyu",
"Zhijian Shao",
"Yuchen Shi",
"Shuwen Zhang",
"Chaofan Qiu",
"Linjie Che"
],
"upvotes": 10,
"arxiv_id": "2607.20911",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Tencent",
"project_page": "https://workbuddybench.com/index.html",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}05GraphVid: Interactive Graph-Controllable Video GenerationControllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple{"url":"https://arxiv.org/abs/2607.21580","tldr":"Controllable video generation …
EVENT. cmrylpjnID. cmrylpjn88nc5kh0c1612pq5sSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21580",
"tldr": "Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movement. In practice, trajectory-based control often requires users to draw accurate tracks for multiple objects, which scales poorly with scene complexity and becomes ambiguous under occlusion or overlap. To enable flexible yet precise multi-subject control, we introduce GraphVid, a graph-conditioned i",
"title": "GraphVid: Interactive Graph-Controllable Video Generation",
"authors": [
"Vedant Shah",
"Onkar Susladkar",
"Tushar Prakash",
"Kiet Nguyen",
"Tianjio Yu",
"Adheesh Juvekar",
"Muntasir Waheed",
"Ismini Lourentzou"
],
"upvotes": 1,
"arxiv_id": "2607.21580",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "PLAN Lab @University of Illinois Urbana-Champaign",
"project_page": "https://plan-lab.github.io/projects/graphvid",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}06AREX: Towards a Recursively Self-Improving Agent for Deep ResearchDeep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do mo{"url":"https://arxiv.org/abs/2607.21461","tldr":"Deep research requires agents …
EVENT. cmrylpixID. cmrylpixn8nc3kh0cbesa33ltSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21461",
"tldr": "Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce ARE",
"title": "AREX: Towards a Recursively Self-Improving Agent for Deep Research",
"authors": [
"Shuqi Lu",
"Chaofan Li",
"Kun Luo",
"Zhang Zhang",
"Hui Wang",
"Hongwang Xiao",
"Zheng Liu",
"Lei Xiong",
"Jiahao Wang",
"Sen Wang",
"Xiyan Jiang",
"Wanli Li"
],
"upvotes": 46,
"arxiv_id": "2607.21461",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Beijing Academy of Artificial Intelligence",
"project_page": "https://vectorspacelab.github.io/arex-model/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}07ReferTrack: Referring Then Tracking for Embodied Visual TrackingEmbodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning ofte{"url":"https://arxiv.org/abs/2607.20061","tldr":"Embodied visual tracking (EVT)…
EVENT. cmrylpi8ID. cmrylpi868nc1kh0ct5j9imvhSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20061",
"tldr": "Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking para",
"title": "ReferTrack: Referring Then Tracking for Embodied Visual Tracking",
"authors": [
"Hanjing Ye",
"Tianle Zeng",
"Jiazhao Zhang",
"Shaoan Wang",
"Zibo Zhang",
"Weisi Situ",
"Yuchen Zhou",
"Yonggen Ling",
"Hong Zhang"
],
"upvotes": 18,
"arxiv_id": "2607.20061",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Tencent",
"project_page": "https://medlartea.github.io/referTrack/",
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}08Robostral NavigateDeploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployme{"url":"https://arxiv.org/abs/2607.20785","tldr":"Deploying navigation systems a…
EVENT. cmrylphiID. cmrylphip8nbzkh0cfl2v18meSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20785",
"tldr": "Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently. Yet, today's best systems depend on depth sensors, multi-camera rigs, or pre-built maps, limiting the hardware they support and increasing deployment cost. We introduce Robostral Navigate, an 8B vision-language model built around this scalability objective. The model consumes only a stream of monocular RGB images - the most ubiquitous sensor acr",
"title": "Robostral Navigate",
"authors": [
"Arjun Majumdar",
"Avinash Sooriyarachchi",
"Benjamin Tibi",
"Chris Bamford",
"Elliot Chane-Sane",
"Guillaume Lample",
"Khyathi Raghavi Chandu",
"Ludovic Ho Fuh",
"Mathieu Poiree",
"Olivier Duchenne",
"Rosalie Millner",
"Srijan Mishra"
],
"upvotes": 3,
"arxiv_id": "2607.20785",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Mistral AI_",
"project_page": "https://mistral.ai/news/robostral-navigate/",
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}09Visual Contrastive Self-DistillationOn-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods cr{"url":"https://arxiv.org/abs/2607.21556","tldr":"On-policy self-distillation (O…
EVENT. cmrylpgtID. cmrylpgt78nbxkh0cc338mh94SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21556",
"tldr": "On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymmetric information between teacher and student to ensure that the self-teacher provides a stronger learning signal than the student. Existing methods create this asymmetry either through privileged answers or visual evidence. We ask whether both can be removed, yielding a simpler form of OPSD driven purely by input conditioning. For this purpose, we ",
"title": "Visual Contrastive Self-Distillation",
"authors": [
"Yijun Liang",
"Yunjie Tian",
"Yijiang Li",
"Yuqi Jia",
"Furong Huang",
"Tianyi Zhou",
"Di Fu"
],
"upvotes": 32,
"arxiv_id": "2607.21556",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/66720ab819bebc69b5b93685/DAWO1NIJw1bQwXlnpqvAw.png",
"ai_summary": null,
"ai_keywords": [],
"organization": "University of Maryland College Park",
"project_page": "https://joliang17.github.io/VisualCSD/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}10Recurrent Sinusoidal INRs for Efficient High-Fidelity RepresentationWe study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectr{"url":"https://arxiv.org/abs/2607.21485","tldr":"We study sinusoidal recurrence…
EVENT. cmrylpg1ID. cmrylpg1x8nbvkh0ch1pdlf2wSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21485",
"tldr": "We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectral support. We realize this principle with a shared sinusoidal block that iteratively refines the latent representation. We empirically validate the resulting spectral behavior against feed-forward IN",
"title": "Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation",
"authors": [
"Hyunmin Cho",
"Jaejun Yoo",
"Kyong Hwan Jin"
],
"upvotes": 5,
"arxiv_id": "2607.21485",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://hyeon-cho.github.io/Harmonic-line-Spectrum/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}showing 1–10 of 498older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above