arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 462 events in this window (598 total on topic). adjust the range or clear it with ALL.
range
01HPD-Parsing: Hierarchical Parallel Document ParsingEfficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autore{"url":"https://arxiv.org/abs/2607.18839","tldr":"Efficient teamwork typically c…
EVENT. cmrvqq4kID. cmrvqq4kt7v1tkh0cef38krd9SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.18839",
"tldr": "Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autoregressive trajectory, creating a sequential bottleneck that grows with document length. Such full-page sequential generation overlooks a key property of document parsing: layout must be analyzed global",
"title": "HPD-Parsing: Hierarchical Parallel Document Parsing",
"authors": [
"Shu Wei",
"Jingjing Wu",
"Lingshu Zhang",
"Qunyi Xie",
"Hao Zou",
"Le Xiang",
"Xu Fan",
"Yangliu Xu",
"Manhui Lin",
"Xiaolong Ma",
"Cheng Cui",
"Tengyu Du"
],
"upvotes": 5,
"arxiv_id": "2607.18839",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "PaddlePaddle",
"project_page": null,
"published_at": "2026-07-21T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}02ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUWe present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven {"url":"https://arxiv.org/abs/2607.19191","tldr":"We present ABot-World-0, an ac…
EVENT. cmrvqq40ID. cmrvqq40e7v1rkh0cdttu9i4iSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.19191",
"tldr": "We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a ",
"title": "ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU",
"authors": [
"Fan Jiang",
"Zhaoxu Sun",
"Mengchao Wang",
"Ziyu Zhu",
"Chiyu Wang",
"Yunpeng Zhang",
"Wenlin Liu",
"Yun Wang",
"Xue Zheng",
"Rui Sun",
"Junfeng Ni",
"Hongyu Pan"
],
"upvotes": 5,
"arxiv_id": "2607.19191",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/ew4mZ_fcuneg7qXzDabcH.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-21T15:26:50.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}03ISO: An RLVR-Native Optimization StackReinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this mis{"url":"https://arxiv.org/abs/2607.19331","tldr":"Reinforcement learning with ve…
EVENT. cmrvqq3fID. cmrvqq3fa7v1pkh0csoyloutiSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.19331",
"tldr": "Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associa",
"title": "ISO: An RLVR-Native Optimization Stack",
"authors": [
"Hanqing Zhu",
"Wenyan Cong",
"Zhizhou Sha",
"Sagnik Mukherjee",
"Xinyuan Song",
"David González-Martínez",
"Xiaoxia Wu",
"Yuandong Tian",
"Shiwei Liu",
"David Z. Pan",
"Zhangyang \"Atlas\" Wang"
],
"upvotes": 1,
"arxiv_id": "2607.19331",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/63884e0a3143f4706312f8eb/7Z2exJEaUm2MPrrsMsxNV.png",
"ai_summary": null,
"ai_keywords": [],
"organization": "University of Texas at Austin",
"project_page": "https://iso-rlvr.github.io/",
"published_at": "2026-07-21T17:51:36.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}04Text Template Tokens Are Implicit Semantic Registers in Diffusion TransformersText-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions{"url":"https://arxiv.org/abs/2607.19139","tldr":"Text-to-image diffusion transf…
EVENT. cmrvqq2uID. cmrvqq2uq7v1nkh0cbll5typ6SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.19139",
"tldr": "Text-to-image diffusion transformers (DiTs) jointly process text and image tokens, yet their internal computation during denoising remains poorly understood. We introduce a causal interpretability framework for modern large-scale DiTs that combines attention decomposition with targeted interventions across token spans, heads, and layers. Using it to separate prompt-content tokens from structural template tokens, we find that the structural tokens carry little prompt-specific information at the e",
"title": "Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers",
"authors": [
"Maohua Li",
"Qirui Li",
"Yanke Zhou",
"Yiduo Li",
"Zhaosheng Chi",
"Chao Xu",
"Cuifeng Shen",
"Yixuan Xu",
"Hanlin Tang",
"Kan Liu",
"Tao Lan",
"Lin Qu"
],
"upvotes": 37,
"arxiv_id": "2607.19139",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "RTP-LLM",
"project_page": null,
"published_at": "2026-07-21T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}05AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM AgentsLLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an op{"url":"https://arxiv.org/abs/2607.18754","tldr":"LLM agent failures are difficu…
EVENT. cmrvqq2aID. cmrvqq2a57v1lkh0cwljnrx0pSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.18754",
"tldr": "LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support for identifying the root cause or translating diagnosis into recovery. We present AgentDebugX, an open-source debugging framework that organizes debugging as a closed loop of Detect, Attribute, Recover, and Rerun. At its core, DeepDebug performs multi-turn root-cause diagnosis through global traject",
"title": "AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents",
"authors": [
"Kunlun Zhu",
"Xuyan Ye",
"Zhiguang Han",
"Yuchen Zhao",
"Bingxuan Li",
"Weijia Zhang",
"Muxin Tian",
"Xiangru Tang",
"Pan Lu",
"James Zou",
"Jiaxuan You",
"Heng Ji"
],
"upvotes": 9,
"arxiv_id": "2607.18754",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "University of Illinois at Urbana-Champaign",
"project_page": "https://www.agentdebugx.com/",
"published_at": "2026-07-21T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}06Generative World Renderer at the Speed of PlayGenerative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the underlying world dynamics. This demonstr{"url":"https://arxiv.org/abs/2607.18703","tldr":"Generative world renderer Alay…
EVENT. cmrvqq1pID. cmrvqq1pk7v1jkh0ccsd3u2azSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.18703",
"tldr": "Generative world renderer AlayaRenderer receives structured world states exported from physics engines and synthesizes RGB frames. Unlike models that generate frames from text/control-hints prompts, AlayaRenderer preserves scene structure without altering the underlying world dynamics. This demonstrates an alternative path toward interactive world modeling and user-controllable play. However, the original AlayaRenderer is too computationally expensive for real-time deployment. This technical rep",
"title": "Generative World Renderer at the Speed of Play",
"authors": [
"Guixu Lin",
"Zheng-Hui Huang",
"Siqi Yang",
"Ming-Hsuan Yang",
"Kaipeng Zhang",
"Zhixiang Wang"
],
"upvotes": 39,
"arxiv_id": "2607.18703",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6307a98795b2ab342fec0cf7/ig-BJzrK-5OFq4Yj_Lsf2.png",
"ai_summary": null,
"ai_keywords": [],
"organization": "Alaya Lab",
"project_page": "https://alaya-renderer-flash.alayalab.ai/",
"published_at": "2026-07-21T04:54:24.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}07EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust CalibrationTeaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the int{"url":"https://arxiv.org/abs/2607.18529","tldr":"Teaching videos are becoming a…
EVENT. cmrvqq14ID. cmrvqq14w7v1hkh0cz97iz4t6SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.18529",
"tldr": "Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting because teaching quality depends on multimodal evidence and should be evaluated with respect to the intended learner rather than as a universal property. We present EduPanel, a rubric-grounded, learner-conditioned LLM judge that decomposes evaluation across specialized agents to produce interpretable a",
"title": "EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration",
"authors": [
"Jia-Kai Dong",
"Yi-Cheng Lin",
"Hung-yi Lee"
],
"upvotes": 1,
"arxiv_id": "2607.18529",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "台灣大學",
"project_page": null,
"published_at": "2026-07-20T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}08ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric VideoEgocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desi{"url":"https://arxiv.org/abs/2607.17790","tldr":"Egocentric devices, such as we…
EVENT. cmrv1032ID. cmrv1032v7o03kh0coduifm9eSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.17790",
"tldr": "Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a human viewer and the surrounding environment. A holistic and efficient multimodal model capable of reconstructing this 4D representation is therefore highly desirable. However, existing approaches often rely on auxiliary inputs such as pre-computed camera trajectories, treat scene perception and human ego-motion modeling as separate problems despite their str",
"title": "ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video",
"authors": [
"Xiaozhong Lyu",
"Gen Li",
"Zhiyin Qian",
"Xucong Zhang",
"Marc Pollefeys",
"Siyu Tang"
],
"upvotes": 2,
"arxiv_id": "2607.17790",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/673e46a74dca9bce3141b78b/FReRnN47ukLHeoCo2ZYWH.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "Computer Vision and Learning Group, ETH Zürich",
"project_page": "https://reviv4d.github.io/",
"published_at": "2026-07-20T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-21T00:00:00.000Z"
}09Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted EscalationMulti-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model{"url":"https://arxiv.org/abs/2607.15434","tldr":"Multi-agent systems routinely …
EVENT. cmrv102jID. cmrv102j57o01kh0c5dclu6yxSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.15434",
"tldr": "Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably decline",
"title": "Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation",
"authors": [
"Jasmine Brazilek",
"Maheep Chaudhary",
"Zoe Lu",
"Miles Tidmarsh"
],
"upvotes": 2,
"arxiv_id": "2607.15434",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-20T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-21T00:00:00.000Z"
}10FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality MimicryIn line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated proc{"url":"https://arxiv.org/abs/2607.18227","tldr":"In line with the prevailing di…
EVENT. cmruo5r1ID. cmruo5r1s7kh5kh0ctd0r7137SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.18227",
"tldr": "In line with the prevailing direction of vision research, we explore the integration of both generation and editing capabilities for video and image modalities within a single model. Current approaches to collecting video editing data typically depend on labour-intensive, time-consuming curated procedures--involving object mask annotation, the use of error-introducing pair synthesis via I2V model and ControlNet-like guidance, and VLM-based quality filtering or refinement--and demonstrate limited",
"title": "FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry",
"authors": [
"Dingyun Zhang",
"Lixue Gong",
"Wei Liu"
],
"upvotes": 13,
"arxiv_id": "2607.18227",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "ByteDance Seed",
"project_page": null,
"published_at": "2026-07-20T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-21T00:00:00.000Z"
}showing 1–10 of 462older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above