arXiv Papers

topic · knowledge/papers-arxiv
DOC.
knowledge/papers-arxiv
REV.
598 evt
DATE.
29-MAY-2026
SCOPE.
custom
§01

about

Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 480 events in this window (598 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object DetectionThis work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotati{"url":"https://arxiv.org/abs/2607.19942","tldr":"This work introduces G-MAD, an…
EVENT. cmrx66j7ID. cmrx66j7989axkh0c071vm59lSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19942",
  "tldr": "This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-le",
  "title": "G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection",
  "authors": [
    "Yechan Kim",
    "JongHyun Park",
    "Dongho Yoon",
    "Namhoon Jung",
    "Moongu Jeon"
  ],
  "upvotes": 0,
  "arxiv_id": "2607.19942",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
02SLPO: Scaling Latent Reasoning via a Surrogate PolicyReinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead c{"url":"https://arxiv.org/abs/2607.19691","tldr":"Reinforcement learning with ve…
EVENT. cmrx66inID. cmrx66iny89avkh0cx27i8d4ySRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19691",
  "tldr": "Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, whil",
  "title": "SLPO: Scaling Latent Reasoning via a Surrogate Policy",
  "authors": [
    "Runyang You",
    "Zhiyuan Liu",
    "Yongqi Li",
    "Wenjie Li"
  ],
  "upvotes": 1,
  "arxiv_id": "2607.19691",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
03SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPODFull-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training syst{"url":"https://arxiv.org/abs/2607.20145","tldr":"Full-parameter post-training o…
EVENT. cmrx66i4ID. cmrx66i4o89atkh0czuz15c11SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.20145",
  "tldr": "Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hi",
  "title": "SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD",
  "authors": [
    "Dongfang Li",
    "Xiaodong Luo",
    "Ruoyu Sun",
    "Xuhui Chen",
    "Linyuan Qiu",
    "Jian Meng",
    "Zhengxuan Lu",
    "Yiting Wang",
    "Yucheng Xie",
    "Tao Guo",
    "Tianxiang Fang",
    "Jing Li"
  ],
  "upvotes": 29,
  "arxiv_id": "2607.20145",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
04Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and RankingAs large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document intera{"url":"https://arxiv.org/abs/2607.19747","tldr":"As large language models and A…
EVENT. cmrx66hkID. cmrx66hkm89arkh0cb3esf939SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19747",
  "tldr": "As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework",
  "title": "Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking",
  "authors": [
    "Kailin Jiang",
    "Lei Liu",
    "Jian Xi",
    "Hui Xu",
    "Junlin Liu",
    "Baochen Fu",
    "Shaoqing Ren",
    "Bin Li",
    "Vichwang",
    "Yu Lu",
    "Haibo Shi"
  ],
  "upvotes": 4,
  "arxiv_id": "2607.19747",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://rubric4setwise.github.io/",
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
05Self Gradient Forcing: Native Long Video ExtrapolationRecent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as{"url":"https://arxiv.org/abs/2607.20368","tldr":"Recent autoregressive video di…
EVENT. cmrx66h1ID. cmrx66h1189apkh0cnkd3hc0gSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.20368",
  "tldr": "Recent autoregressive video diffusion methods are increasingly built upon Self Forcing, where the student is trained on histories produced by its own rollout rather than ground-truth video contexts. This reduces exposure bias, but the historical key-value cache is still used by future frames only as frozen rollout state. As a result, future losses cannot supervise how earlier generated latents should be written into more useful keys and values for later video-latent generation. We call this the ",
  "title": "Self Gradient Forcing: Native Long Video Extrapolation",
  "authors": [
    "Junhao Zhuang",
    "Shiyi Zhang",
    "Yuxuan Bian",
    "Yaowei Li",
    "Yawen Luo",
    "Yijun Liu",
    "Weiyang Jin",
    "Songchun Zhang",
    "Xianglong He",
    "Xuying Zhang",
    "Haoran Li",
    "Haoyang Huang"
  ],
  "upvotes": 19,
  "arxiv_id": "2607.20368",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://zhuang2002.github.io/SelfGradientForcing/",
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
06Trace: A Taxonomy-Guided Environment for Multidomain Visual ReasoningReinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduce Trace, a taxonomy-{"url":"https://arxiv.org/abs/2607.19790","tldr":"Reinforcement learning with ve…
EVENT. cmrx66ghID. cmrx66ghs89ankh0c0ky8r9nlSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.19790",
  "tldr": "Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data that are simultaneously broad, exactly verifiable, and reproducible. We introduce Trace, a taxonomy-guided environment for multidomain visual reasoning. Trace factorizes task construction into a scene grammar and an executable task program, separating visual realization from answer computation. A sh",
  "title": "Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning",
  "authors": [
    "Md Tanvirul Alam"
  ],
  "upvotes": 2,
  "arxiv_id": "2607.19790",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/69bc258a4122dc633a350fd7/Y8YZ_mWsa_7YdY5WA5BOs.png",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": "https://maveryn.github.io/trace/",
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
07Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation ExplanationsNatural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim{"url":"https://arxiv.org/abs/2607.20379","tldr":"Natural-language autoencoders …
EVENT. cmrx66fxID. cmrx66fx889alkh0cv2qfrwwbSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.20379",
  "tldr": "Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are reconst",
  "title": "Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations",
  "authors": [
    "Hiskias Dingeto"
  ],
  "upvotes": 0,
  "arxiv_id": "2607.20379",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}
08Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language ModelLarge language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: c{"url":"https://arxiv.org/abs/2607.20058","tldr":"Large language models can answ…
EVENT. cmrwtcjzID. cmrwtcjz68623kh0ch9yu69f8SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.20058",
  "tldr": "Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering a",
  "title": "Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model",
  "authors": [
    "Markus J. Buehler"
  ],
  "upvotes": 0,
  "arxiv_id": "2607.20058",
  "media_url": "https://cdn-uploads.huggingface.co/production/uploads/623ce1c6b66fedf374859fe7/bzxK3-zRnXW_vmx0WiMSJ.png",
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "LAMM: MIT Laboratory for Atomistic and Molecular Mechanics",
  "project_page": null,
  "published_at": "2026-07-22T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
09NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMsScaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting{"url":"https://arxiv.org/abs/2607.14186","tldr":"Scaling executable agent train…
EVENT. cmrvqqaaID. cmrvqqaaa7v2dkh0cwdbu5cs9SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.14186",
  "tldr": "Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage requires manual substrate engineering, each new domain demands a bespoke pipeline, and the resulting task distributions often reflect substrate biases rather than real-world demand. We introduce NexForge, a requirement-driven framework that takes high-level capability requirements as input and synth",
  "title": "NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs",
  "authors": [
    "Jiarong Zhao",
    "Zhikai Lei",
    "Zhiheng Xi",
    "Rui Zheng",
    "Hang Yan",
    "Jie Zhou",
    "Qin Chen",
    "Liang He"
  ],
  "upvotes": 4,
  "arxiv_id": "2607.14186",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": "Nex AGI",
  "project_page": "https://nex.sii.edu.cn/",
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-21T00:00:00.000Z"
}
10H^2SD: Hybrid Hindsight Self-DistillationReinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire trajectory, resulting in sparse sup{"url":"https://arxiv.org/abs/2607.18955","tldr":"Reinforcement learning with ve…
EVENT. cmrvqq9pID. cmrvqq9pq7v2bkh0chvyigyreSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.18955",
  "tldr": "Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models on tasks such as mathematical reasoning and code generation. However, most RLVR methods assign a scalar outcome reward to an entire trajectory, resulting in sparse supervision and limited token-level credit assignment. On-policy distillation (OPD) provides denser supervision by distilling token-level distributions from a stronger teacher model, but requires an addi",
  "title": "H^2SD: Hybrid Hindsight Self-Distillation",
  "authors": [
    "Qiye Cai",
    "Yichuan Ma",
    "Linyang Li",
    "Peiji Li",
    "Yongkang Chen",
    "Qipeng Guo",
    "Yicheng Zou",
    "Tao Gui",
    "Xiaocheng Feng",
    "Bing Qin"
  ],
  "upvotes": 1,
  "arxiv_id": "2607.18955",
  "media_url": null,
  "ai_summary": null,
  "ai_keywords": [],
  "organization": null,
  "project_page": null,
  "published_at": "2026-07-21T00:00:00.000Z",
  "submitted_on_daily_at": "2026-07-22T00:00:00.000Z"
}
showing 1–10 of 480older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above