arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 516 events in this window (596 total on topic). adjust the range or clear it with ALL.
range
01Rethinking Classifier-Free Guidance in On-Policy Diffusion DistillationOn-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally ex{"url":"https://arxiv.org/abs/2607.24731","tldr":"On-policy distillation (OPD) a…
EVENT. cms4bju1ID. cms4bju1ja2zvkh0ctvzcwzlcSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24731",
"tldr": "On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negat",
"title": "Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation",
"authors": [
"Bingnan Li",
"Haozhe Wang",
"Haozhong Xiong",
"Fangtai Wu",
"Jinpeng Yu",
"Yang Shi",
"Jiaming Liu",
"Ruihua Huang"
],
"upvotes": 16,
"arxiv_id": "2607.24731",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://rethinking-cfg-opd.github.io",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}02IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic LanguagesLarge Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We pres{"url":"https://arxiv.org/abs/2607.23242","tldr":"Large Language Models (LLMs) h…
EVENT. cms4bjthID. cms4bjthha2zrkh0c99ub4skySRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.23242",
"tldr": "Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic ",
"title": "IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages",
"authors": [
"Sahil Deepak Gawande",
"Mayank Singh"
],
"upvotes": 1,
"arxiv_id": "2607.23242",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Lingo Research Group",
"project_page": null,
"published_at": "2026-07-25T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}03Evidence Attribution in Visual Document Understanding without Coordinates or Region LabelsReliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document{"url":"https://arxiv.org/abs/2607.24651","tldr":"Reliable visual document under…
EVENT. cms4bjsxID. cms4bjsxea2zpkh0c5j19f76jSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24651",
"tldr": "Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investiga",
"title": "Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels",
"authors": [
"Zhuchenyang Liu",
"Yao Zhang",
"Yu Xiao"
],
"upvotes": 1,
"arxiv_id": "2607.24651",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}04The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic DistillationMulti-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address {"url":"https://arxiv.org/abs/2607.24720","tldr":"Multi-turn long-horizon planni…
EVENT. cms4bjsdID. cms4bjsd8a2znkh0cp9clk94xSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24720",
"tldr": "Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning abilit",
"title": "The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation",
"authors": [
"Tianyi Men",
"Zhuoran Jin",
"Kang Liu",
"Jun Zhao"
],
"upvotes": 12,
"arxiv_id": "2607.24720",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Chinese Academic of Science Institute of Automation",
"project_page": "https://quester-one.github.io/PlanPhysWebsite/",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}05Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed SamplingThe rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single{"url":"https://arxiv.org/abs/2607.23518","tldr":"The rapid evolution of generat…
EVENT. cms4bjrsID. cms4bjrsla2zlkh0coud585whSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.23518",
"tldr": "The rapid evolution of generative models has unlocked new potentials in protein binder design, a pivotal task in structural biology, by facilitating end-to-end generation via joint sequence-structure modeling or hallucination. However, existing approaches are predominantly implemented under a single-target, single-state assumption, limiting their ability to model multi-target or multi-state interactions required for advanced function-oriented protein design. Here, we introduce Chamaileon, which ",
"title": "Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling",
"authors": [
"Hengyuan Cao",
"Shizhuo Cheng",
"Mingxuan Liu",
"Weicheng Huang",
"Yunhong Lu",
"Chenxi Cai",
"Yan Zhang",
"Min Zhang"
],
"upvotes": 6,
"arxiv_id": "2607.23518",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Zhejiang University",
"project_page": null,
"published_at": "2026-07-26T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}06Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention SparsificationDiffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify a{"url":"https://arxiv.org/abs/2607.24027","tldr":"Diffusion transformers are ess…
EVENT. cms4bjr7ID. cms4bjr7xa2zjkh0cl7nxs82eSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24027",
"tldr": "Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this bottleneck by computing only selected key-value blocks, yet existing methods struggle to sparsify attention both efficiently and accurately for two reasons: (1) Rigid, unpredictable, and costly routing: selecting a fixed fraction of top-ranked blocks by proxy score imposes fixed budgets, whereas re",
"title": "Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification",
"authors": [
"Haopeng Li",
"Yitong Li",
"Junsong Chen",
"Tian Ye",
"Haozhe Liu",
"Jincheng Yu",
"Duomin Wang",
"Ruihua Zhang",
"Zeke Xie",
"Enze Xie",
"Song Han"
],
"upvotes": 9,
"arxiv_id": "2607.24027",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "http://nvlabs.github.io/Sana/Sol-Attn/",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}07Spectral Prior for Reducing Exposure Bias in Diffusion ModelsDiffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies between training and inference, which can be interpreted as frequency-dependent SNR error. Crucially, the direction of th{"url":"https://arxiv.org/abs/2607.22091","tldr":"Diffusion models typically suf…
EVENT. cms38xqyID. cms38xqy99su9kh0cokuwpq4lSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.22091",
"tldr": "Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequency-dependent discrepancies between training and inference, which can be interpreted as frequency-dependent SNR error. Crucially, the direction of this mismatch varies across models and timesteps, indicating that fixed correction rules do not generalize. We propose Spectral Alignment (SPA), a lightweight, guidance-based method that calibrates the ",
"title": "Spectral Prior for Reducing Exposure Bias in Diffusion Models",
"authors": [
"Yuya Kobayashi",
"Masato Ishii",
"Yuhta Takida",
"Takashi Shibuya",
"Yuki Mitsufuji"
],
"upvotes": 3,
"arxiv_id": "2607.22091",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Sony",
"project_page": null,
"published_at": "2026-07-24T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}08Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving SkillsLLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound methods obtain precise feedback but confine learning to narro{"url":"https://arxiv.org/abs/2607.22529","tldr":"LLM training is shifting from …
EVENT. cms2w1uxID. cms2w1ux89pjhkh0c685ud99wSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.22529",
"tldr": "LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound methods obtain precise feedback but confine learning to narrow domains, while open-ended self-generation broadens the task space but lacks reliable verification, allowing misleading rewards to pollute the training loop. We identify agent skills as a powerful mi",
"title": "Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills",
"authors": [
"Siyuan Huang",
"Pengyu Cheng",
"Haotian Liu",
"Tao Chen",
"Yihao Liu",
"Jingwei Ni",
"Shijie Zhou",
"Ziyi Yang",
"Gangwei Jiang",
"Mengyu Zhou",
"Yu Cheng",
"Xiaoxi Jiang"
],
"upvotes": 14,
"arxiv_id": "2607.22529",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "QwenBusinessUnit-Edu",
"project_page": null,
"published_at": "2026-07-24T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}09SceneActBench: Can Agents Act on the 3D Scenes They See?Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for{"url":"https://arxiv.org/abs/2607.22393","tldr":"Vision-language model (VLM) ag…
EVENT. cms2w1ucID. cms2w1uc39pjfkh0czlwgqep5SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.22393",
"tldr": "Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent action on complete multi-object 3D scenes under evaluated. We present SceneActBench, a benchmark for visually conditioned action across five 3D tasks under a unified agent-environment loop. Given PNG images or sampled video frames and, where applicable, supplied 3D assets, an agent acts on a 3D envi",
"title": "SceneActBench: Can Agents Act on the 3D Scenes They See?",
"authors": [
"Yifei Zhao",
"Xiangxin Zhou",
"Wenhao Yang",
"Jiaqi Tang",
"Pu Jian",
"Huanjin Yao",
"Jiarui Yao",
"Haowei Lin",
"Chunchao Guo",
"Zhuo Chen",
"Wenkai Lyu",
"Jianzhu Ma"
],
"upvotes": 3,
"arxiv_id": "2607.22393",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://feinaldo2.github.io/sceneactbench-project-page/",
"published_at": "2026-07-24T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}10Scaling Native Multimodal Pre-Training From ScratchAlthough large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby ach{"url":"https://arxiv.org/abs/2607.22043","tldr":"Although large language models…
EVENT. cms2w1swID. cms2w1swa9pjbkh0c3szxpqhxSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.22043",
"tldr": "Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain ",
"title": "Scaling Native Multimodal Pre-Training From Scratch",
"authors": [
"Haoyuan Wu",
"Aoqi Wu",
"Hai Wang",
"Jiajia Wu",
"Jinxiang Ou",
"Bei Yu"
],
"upvotes": 7,
"arxiv_id": "2607.22043",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Tencent Hunyuan",
"project_page": null,
"published_at": "2026-07-24T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}showing 1–10 of 516older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above