arXiv Papers

topic · knowledge/papers-arxiv
DOC.
knowledge/papers-arxiv
REV.
596 evt
DATE.
29-MAY-2026
SCOPE.
custom
§01

about

Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 525 events in this window (596 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic SearchAgentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced pr{"url":"https://arxiv.org/abs/2607.24280","tldr":"Agentic search enables large l…
EVENT. cms4bjz2ID. cms4bjz2ga30dkh0cvhqnz4jhSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24280",
  "tldr": "Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded",
  "title": "From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search",
  "authors": [
    "Junlin Liu",
    "Jiangwang Chen",
    "Zixin Song",
    "Shuaiyu Zhou",
    "Chunji Lv",
    "Hank Wu",
    "Kailin Jiang",
    "Jinyang Wu",
    "Bohan Yu",
    "Chenxi Zhou"
  ],
  "upvotes": 42,
  "arxiv_id": "2607.24280",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-27T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
02DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data RecipesWhile data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion f{"url":"https://arxiv.org/abs/2607.24516","tldr":"While data curation for Vision…
EVENT. cms4bjyiID. cms4bjyifa30bkh0cbv7s4qhlSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24516",
  "tldr": "While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains largely heuristic: practitioners stack datasets that pass quality filters, set cross-domain ratios by intuition, and lack a principled, attributable criterion for admitting new data, while frontier recipes remain undisclosed. We formulate data construction as a systematic mixture-optimization problem and turn it into a reproducible engineering discipline by ",
  "title": "DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes",
  "authors": [
    "Jiahao Xie",
    "Zhongbin Guo",
    "Qianle Wang",
    "Ruiqi Lu",
    "Dongling Xiao",
    "Wanxuan Sun",
    "Cheng Yang"
  ],
  "upvotes": 5,
  "arxiv_id": "2607.24516",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-27T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
03JarvisHub: An Open Harness for Canvas-Native Multimodal Creative AgentsCreative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isola{"url":"https://arxiv.org/abs/2607.23588","tldr":"Creative AI is moving from sin…
EVENT. cms4bjxyID. cms4bjxyfa309kh0c4dd9jqceSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.23588",
  "tldr": "Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an ev",
  "title": "JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents",
  "authors": [
    "Yunlong Lin",
    "Zixu Lin",
    "Zhaohu Xing",
    "Biqiang Li",
    "Chenxin Li",
    "Haonan Wang",
    "Haitao Wu",
    "Hengyu Liu",
    "Jianghai Chen",
    "Kaituo Feng",
    "Kaixin Li",
    "Shawn Chen"
  ],
  "upvotes": 56,
  "arxiv_id": "2607.23588",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/64ecb174f22081b4ac7ca397/ClWfgSovGMxph7qyWQeSJ.mp4",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://www.jarvishub.site/",
  "published_at": "2026-07-26T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
04Kimi K3: Open Frontier IntelligenceWe introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model d{"url":"https://arxiv.org/abs/2607.24653","tldr":"We introduce Kimi K3, a 2.8T p…
EVENT. cms4bjxeID. cms4bjxeja307kh0cqh0frt0xSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24653",
  "tldr": "We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in o",
  "title": "Kimi K3: Open Frontier Intelligence",
  "authors": [
    "Kimi Team",
    "Tongtong Bai",
    "Yifan Bai",
    "Yiping Bao",
    "M. C.",
    "Jianfeng Cai",
    "Xinyuan Cai",
    "Peizhou Cao",
    "Yuxuan Cao",
    "Ziwei Chai",
    "Y. Charles",
    "H. S. Che"
  ],
  "upvotes": 127,
  "arxiv_id": "2607.24653",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Moonshot AI",
  "project_page": "https://www.kimi.com/blog/kimi-k3",
  "published_at": "2026-07-27T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
05FilmBench: A Film-Grade Benchmark for Cinematic Video GenerationProgress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies rem{"url":"https://arxiv.org/abs/2607.24241","tldr":"Progress in video generation k…
EVENT. cms4bjwuID. cms4bjwufa305kh0c0vwoh8l9SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24241",
  "tldr": "Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score them with untrained, generic multimodal models. More fundamentally, their evaluation taxonomies remain rudimentary (overall visual quality, coarse text alignment and temporal smoothness) rather than the professional Cinematic Language criteria by which films are actually made and judged, so they as",
  "title": "FilmBench: A Film-Grade Benchmark for Cinematic Video Generation",
  "authors": [
    "Shengyi Wang",
    "Niantong Li",
    "Guangzheng Hu",
    "Hong Qi",
    "Fei Ding",
    "Weixu Qiao",
    "Jinlin Wang",
    "Xiaotong Lv",
    "Peng Han",
    "Zimeng Li",
    "Fanshu Ding",
    "Yushu Wang"
  ],
  "upvotes": 1,
  "arxiv_id": "2607.24241",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-27T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
06GNM Head: A Generative aNthropometric Model of the human headParametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated im{"url":"https://arxiv.org/abs/2607.23687","tldr":"Parametric models of the human…
EVENT. cms4bjwaID. cms4bjwada303kh0c6p4ij08eSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.23687",
  "tldr": "Parametric models of the human head are essential tools traditionally used in computer vision and graphics for animation, rendering, and reconstruction. More recently, they serve as crucial conditioning signals within generative large vision models, allowing for tight spatial control of generated imagery. However, existing publicly available models are typically limited in anatomical scope, modeling only outer geometry while ignoring intra-oral and ocular structures, and frequently suffer from r",
  "title": "GNM Head: A Generative aNthropometric Model of the human head",
  "authors": [
    "Stylianos Ploumpis",
    "Jan Bednarik",
    "Gaspard Zoss",
    "Ruslan Guseinov",
    "Luca Prasso",
    "Prashanth Chandran",
    "Oliver Boyne",
    "Vasileios Choutas",
    "Timo Bolkart",
    "Daoye Wang",
    "Menglei Chai",
    "Di Qiu"
  ],
  "upvotes": 2,
  "arxiv_id": "2607.23687",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Google",
  "project_page": null,
  "published_at": "2026-07-26T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
07A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, ForeverImproving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an indepen{"url":"https://arxiv.org/abs/2607.23806","tldr":"Improving a language model tod…
EVENT. cms4bjvqID. cms4bjvq8a301kh0cs0tyjbgmSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.23806",
  "tldr": "Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take the opposite path: the model stays frozen, and a persistent memory of verified solutions grows beside it. Once a problem family is solved and has passed an independent verification step that never consults the answer key, every new instance of that family is answered at zero generation tokens, bit-exact, deterministically. Across 180 fresh instances spanning ni",
  "title": "A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever",
  "authors": [
    "Sietse Schelpe"
  ],
  "upvotes": 2,
  "arxiv_id": "2607.23806",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Corbenic",
  "project_page": "https://arxiv.org/abs/2607.23806",
  "published_at": "2026-07-26T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
08dRAE: Representation Autoencoder with Hyper-Spherical CodesIn this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric {"url":"https://arxiv.org/abs/2607.22148","tldr":"In this work, we aim to discre…
EVENT. cms4bjv6ID. cms4bjv60a2zzkh0c61zmbuz9SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.22148",
  "tldr": "In this work, we aim to discretize the high-dimensional visual representations to bridge the gap with language models - a non-trivial challenge, as existing quantization methods suffer from codebook collapse, failing to scale while preserving semantic coherence. We identify the root cause as metric mismatch: standard Euclidean codebook objectives are fundamentally misaligned with the anisotropic geometry of representation space, leading to codebook embeddings with high-variance magnitude scales ",
  "title": "dRAE: Representation Autoencoder with Hyper-Spherical Codes",
  "authors": [
    "Tianren Ma",
    "Lin Long",
    "Chuyan Chen",
    "Mu Zhang",
    "Junbo Zhao",
    "Tong Zhang",
    "Qixiang Ye"
  ],
  "upvotes": 3,
  "arxiv_id": "2607.22148",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Ant Group",
  "project_page": "https://drae-hsq.github.io/",
  "published_at": "2026-07-24T09:47:32.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
09Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language ModelsHistorical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external{"url":"https://arxiv.org/abs/2607.21936","tldr":"Historical documents act as in…
EVENT. cms4bjulID. cms4bjulra2zxkh0c1pinipu8SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.21936",
  "tldr": "Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing restoration methods based on masked language modeling effectively utilize local context, they struggle to restore named entities that require external historical knowledge. To address this limitation, we introduce a novel framework for historical document restoration that leverages large language models with retrieval-augmented generation (RAG). By",
  "title": "Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models",
  "authors": [
    "Gabeen Kim",
    "Kyeongpil Kang"
  ],
  "upvotes": 1,
  "arxiv_id": "2607.21936",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Kangwon National University",
  "project_page": null,
  "published_at": "2026-07-24T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
10Rethinking Classifier-Free Guidance in On-Policy Diffusion DistillationOn-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally ex{"url":"https://arxiv.org/abs/2607.24731","tldr":"On-policy distillation (OPD) a…
EVENT. cms4bju1ID. cms4bju1ja2zvkh0ctvzcwzlcSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24731",
  "tldr": "On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally extend velocity matching to the CFG-composed prediction, directly matching teacher and student guided velocities. We show that this objective is under-identified at the branch level: positive- and negat",
  "title": "Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation",
  "authors": [
    "Bingnan Li",
    "Haozhe Wang",
    "Haozhong Xiong",
    "Fangtai Wu",
    "Jinpeng Yu",
    "Yang Shi",
    "Jiaming Liu",
    "Ruihua Huang"
  ],
  "upvotes": 16,
  "arxiv_id": "2607.24731",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://rethinking-cfg-opd.github.io",
  "published_at": "2026-07-27T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}
showing 1–10 of 525older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above