arXiv Papers

topic · knowledge/papers-arxiv
DOC.
knowledge/papers-arxiv
REV.
598 evt
DATE.
29-MAY-2026
SCOPE.
custom
§01

about

Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 471 events in this window (598 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01H^2SD: Hybrid Hindsight Self-DistillationReinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire trajectory, resulting in sparse sup{"url":"https://arxiv.org/abs/2607.18955","tldr":"Reinforcement learning with ve…
EVENT. cmrvqq9pID. cmrvqq9pq7v2bkh0chvyigyreSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.18955",
  "tldr": "Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire trajectory, resulting in sparse supervision and limited token-level credit assignment. On-policy distillation (OPD) provides denser supervision by distilling token-level distributions from a stronger teacher model, but requires an addi",
  "title": "H^2SD: Hybrid Hindsight Self-Distillation",
  "authors": [
    "Qiye Cai",
    "Yichuan Ma",
    "Linyang Li",
    "Peiji Li",
    "Yongkang Chen",
    "Qipeng Guo",
    "Yicheng Zou",
    "Tao Gui",
    "Xiaocheng Feng",
    "Bing Qin"
  ],
  "upvotes": 1,
  "arxiv_id": "2607.18955",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
02Appearance Pointers -- Multimodal Region Control of Diffusion TransformersControllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest hetero{"url":"https://arxiv.org/abs/2607.19344","tldr":"Controllable image generation …
EVENT. cmrvqq95ID. cmrvqq9507v29kh0cxaaqci9qSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19344",
  "tldr": "Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, object identities, and spatial arrangements that cannot be reliably achieved through text prompting alone. Diffusion Transformers (DiTs) can natively ingest heterogeneous tokens stemming from texts and images, but they lack mechanisms for determining where and how these tokens should influence the output. We introduce appearance pointers, compact tokens that gu",
  "title": "Appearance Pointers -- Multimodal Region Control of Diffusion Transformers",
  "authors": [
    "Rahul Sajnani",
    "Yulia Gryaditskaya",
    "Radomír Měch",
    "Srinath Sridhar",
    "Matheus Gadelha"
  ],
  "upvotes": 0,
  "arxiv_id": "2607.19344",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/R1DoG9CXP0SnrIQdJHyM6.png",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://ivl.cs.brown.edu/research/appearance_pointers.html",
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
03Masked Visual Actions for Unified World ModelingVideo models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these inte{"url":"https://arxiv.org/abs/2607.19343","tldr":"Video models absorb rich prior…
EVENT. cmrvqq8kID. cmrvqq8ko7v27kh0cgoknjnfpSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19343",
  "tldr": "Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrar",
  "title": "Masked Visual Actions for Unified World Modeling",
  "authors": [
    "Hadi Alzayer",
    "Wenlong Huang",
    "Haonan Chen",
    "Christopher Luey",
    "Lvmin Zhang",
    "Maneesh Agrawala",
    "Gordon Wetzstein",
    "Li Fei-Fei",
    "Yilun Du",
    "Jiajun Wu",
    "Jia-Bin Huang"
  ],
  "upvotes": 0,
  "arxiv_id": "2607.19343",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://masked-visual-actions.github.io/",
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
04ConsiSpace: Learning Geometric Consistency Matters for Video Spatial ReasoningVideo spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often f{"url":"https://arxiv.org/abs/2607.17599","tldr":"Video spatial reasoning is ess…
EVENT. cmrvqq80ID. cmrvqq8057v25kh0cjhrsf952SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.17599",
  "tldr": "Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consis",
  "title": "ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning",
  "authors": [
    "Ting Huang",
    "Zhenyu Zhang",
    "Wenyuan Huang",
    "Jian Yang",
    "Hao Tang"
  ],
  "upvotes": 0,
  "arxiv_id": "2607.17599",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://believeht029.github.io/ConsiSpace/",
  "published_at": "2026-07-20T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
05Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and EditingLarge-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a l{"url":"https://arxiv.org/abs/2607.19064","tldr":"Large-scale visual generators …
EVENT. cmrvqq7fID. cmrvqq7fq7v23kh0ctpz0zgmfSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19064",
  "tldr": "Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding w",
  "title": "Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing",
  "authors": [
    "Xinjie Zhang",
    "Peng Zhang",
    "Shicheng Zheng",
    "Jinghao Guo",
    "Zhaoyang Jia",
    "Yifei Shen",
    "Xun Guo",
    "Yuxuan Luo",
    "Jiahao Li",
    "Wenxuan Xie",
    "Fanyi Pu",
    "Xiaoyi Zhang"
  ],
  "upvotes": 33,
  "arxiv_id": "2607.19064",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Microsoft",
  "project_page": "https://microsoft.github.io/Mage/",
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
06SciForma: Structure-Faithful Generation of Scientific DiagramsStructural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire fi{"url":"https://arxiv.org/abs/2607.18091","tldr":"Structural fidelity is essenti…
EVENT. cmrvqq6vID. cmrvqq6ve7v21kh0cg6p7x1lySRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.18091",
  "tldr": "Structural fidelity is essential to scientific methodology diagrams. To communicate research logic, these diagrams must faithfully render components, directional relations, and textual annotations. Since a single error, such as a reversed arrow or an unreadable equation, can invalidate the entire figure, structural fidelity is inherently conjunctive: correctness on one axis cannot compensate for failure on another. Current open-source models fail to satisfy this criterion. Supervised fine-tuning",
  "title": "SciForma: Structure-Faithful Generation of Scientific Diagrams",
  "authors": [
    "Yuxuan Luo",
    "Peng Zhang",
    "Xinjie Zhang",
    "Xun Guo",
    "Zhouhui Lian",
    "Yan Lu"
  ],
  "upvotes": 9,
  "arxiv_id": "2607.18091",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-20T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
07Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement LearningAsynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference{"url":"https://arxiv.org/abs/2607.18722","tldr":"Asynchronous reinforcement lea…
EVENT. cmrvqq6aID. cmrvqq6av7v1zkh0cb6w9fyriSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.18722",
  "tldr": "Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint. As a resu",
  "title": "Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning",
  "authors": [
    "Junyao Yang",
    "Yucheng Shi",
    "Zongxia Li",
    "Zhongzhi Li",
    "Ruhan Wang",
    "Xiangxin Zhou",
    "Kishan Panaganti",
    "Haitao Mi",
    "Leowei Liang"
  ],
  "upvotes": 24,
  "arxiv_id": "2607.18722",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Tencent Hunyuan",
  "project_page": null,
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
08AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical ReportUnlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving v{"url":"https://arxiv.org/abs/2607.18367","tldr":"Unlike conventional video game…
EVENT. cmrvqq5qID. cmrvqq5q97v1xkh0cb9wxb88qSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.18367",
  "tldr": "Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly. It enable us to create customized, explorable, and continuously evolving virtual world from text, an image, or video. Realizing this vision requires four tightly coupled capabilities: interaction, persistent spatiotemporal consistency, stable long-horizon generation, and ef",
  "title": "AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report",
  "authors": [
    "AlayaWorld Team",
    "Kaipeng Zhang",
    "Chuanhao Li",
    "Yifan Zhan",
    "Yongtao Ge",
    "Yuanyang Yin",
    "Jiaming Tan",
    "Kang He",
    "Liaoyuan Fan",
    "Mingliang Zhai",
    "Ruicong Liu",
    "Xiaojie Xu"
  ],
  "upvotes": 30,
  "arxiv_id": "2607.18367",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/65f1713552c38a91e0a445e8/Ov_TJfICPbFCAscHvkW98.mp4",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Alaya Lab",
  "project_page": "https://alaya-lab.github.io/AlayaWorld/",
  "published_at": "2026-07-20T17:15:41.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
09Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenessEvaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it shou{"url":"https://arxiv.org/abs/2607.19322","tldr":"Evaluating the factuality of l…
EVENT. cmrvqq55ID. cmrvqq55u7v1vkh0c8x41ntn3SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19322",
  "tldr": "Evaluating the factuality of long-form generations has focused predominantly on precision, measuring whether the claims a model makes are correct. The dominant decompose-search-verify pipeline catches incorrect claims well but says little about whether a response contains all the information it should. Measuring factual completeness, the missing half of factuality, is harder: it requires enumerating the full set of facts a complete answer should contain, and these facts rarely form a flat list. ",
  "title": "Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness",
  "authors": [
    "Xilun Chen",
    "Zhaleh Feizollahi",
    "Ross Goodwin",
    "Seungwhan Moon",
    "Scott Yih",
    "Pinar Donmez",
    "Babak Damavandi",
    "Luna Dong"
  ],
  "upvotes": 3,
  "arxiv_id": "2607.19322",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "AI at Meta",
  "project_page": null,
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
10HPD-Parsing: Hierarchical Parallel Document ParsingEfficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autore{"url":"https://arxiv.org/abs/2607.18839","tldr":"Efficient teamwork typically c…
EVENT. cmrvqq4kID. cmrvqq4kt7v1tkh0cef38krd9SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.18839",
  "tldr": "Efficient teamwork typically combines global coordination with parallel execution, a principle not yet fully reflected in unified Vision-Language Model (VLM)-based document parsers. Existing unified parsers process an entire page jointly but generate its output through a single token-by-token autoregressive trajectory, creating a sequential bottleneck that grows with document length. Such full-page sequential generation overlooks a key property of document parsing: layout must be analyzed global",
  "title": "HPD-Parsing: Hierarchical Parallel Document Parsing",
  "authors": [
    "Shu Wei",
    "Jingjing Wu",
    "Lingshu Zhang",
    "Qunyi Xie",
    "Hao Zou",
    "Le Xiang",
    "Xu Fan",
    "Yangliu Xu",
    "Manhui Lin",
    "Xiaolong Ma",
    "Cheng Cui",
    "Tengyu Du"
  ],
  "upvotes": 5,
  "arxiv_id": "2607.18839",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "PaddlePaddle",
  "project_page": null,
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
showing 1–10 of 471older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above