arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago26 evt / 1h
showing 10 of 552 events in this window (596 total on topic). adjust the range or clear it with ALL.
range
01Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine DetectionHyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspected before targets are found. This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using {"url":"https://arxiv.org/abs/2607.25310","tldr":"Hyperspectral imaging (HSI) is…
EVENT. cms6tk0nID. cms6tk0n0aqspkh0ch4g5cgwfSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25310",
"tldr": "Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspected before targets are found. This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using spectral angle mapper (SAM), matched filter (MF), adaptive coherence estimator (ACE), and constrained energy minimization (CEM). We compare a ground-measured SVC signature, a fully informed in-scene c",
"title": "Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection",
"authors": [
"Sagar Lekhak",
"Prasanna Reddy Pulakurthi",
"Emmett J. Ientilucci"
],
"upvotes": 1,
"arxiv_id": "2607.25310",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Rochester Institute of Technology",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}02Reinforcement Learning for Code OptimizationRL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small pro{"url":"https://arxiv.org/abs/2607.25970","tldr":"RL for code correctness is now…
EVENT. cms6gov4ID. cms6gov4zan7dkh0c2a3otr91SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25970",
"tldr": "RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small problems in measurement noise, reward sparsity, or GRPO instability overwhelm the signal and make RL fail: generated solutions are barely faster, and more of them can fail. We make execution time learnab",
"title": "Reinforcement Learning for Code Optimization",
"authors": [
"Pierre Chambon",
"Kunhao Zheng",
"Juliette Decugis",
"Benoit Sagot",
"Gabriel Synnaeve"
],
"upvotes": 3,
"arxiv_id": "2607.25970",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "AI at Meta",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}03Pass the Baton: Trajectory-Relayed On-Policy DistillationOn-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable super{"url":"https://arxiv.org/abs/2607.26057","tldr":"On-policy distillation (OPD) g…
EVENT. cms63uyeID. cms63uyesajtjkh0cw0pckox5SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.26057",
"tldr": "On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and con",
"title": "Pass the Baton: Trajectory-Relayed On-Policy Distillation",
"authors": [
"Haolei Xu",
"Xiaowen Xu",
"Haiwen Hong",
"Zixuan Ni",
"Hongxing Li",
"Yiwen Qiu",
"Weiming Lu",
"Yongliang Shen"
],
"upvotes": 16,
"arxiv_id": "2607.26057",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Zhejiang University",
"project_page": "https://zju-real.github.io/Relay-OPD/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}04MODUS: Decoder-Only Any-to-Any Modeling of Diverse ModalitiesAny-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using{"url":"https://arxiv.org/abs/2607.25948","tldr":"Any-to-any models predict any …
EVENT. cms63uxvID. cms63uxvxajthkh0cw33ua0gvSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25948",
"tldr": "Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any",
"title": "MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities",
"authors": [
"Mingqiao Ye",
"Zhaochong An",
"Zhitong Gao",
"Xian Liu",
"François Fleuret",
"Chuan Li",
"Amir Zadeh",
"Serge Belongie",
"Afshin Dehghan",
"Jesse Allardice",
"David Mizrahi",
"Oğuzhan Fatih Kar"
],
"upvotes": 6,
"arxiv_id": "2607.25948",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/648b942c086aad6cd6c1d3a4/xfVnochs_G6gCNx7PdzHO.png",
"ai_summary": null,
"ai_keywords": [],
"organization": "EPFL VILAB",
"project_page": "https://modus-multimodal.epfl.ch/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}05Parallel Decoding Distillation for Fast Image and Video GenerationGeneration in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generat{"url":"https://arxiv.org/abs/2607.26004","tldr":"Generation in video diffusion …
EVENT. cms63uxcID. cms63uxclajtdkh0cpt6dijvaSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.26004",
"tldr": "Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variational score distillation (VSD) and adversarial losses to distill diffusion models into few-step generators. Albeit achieving high-quality video generation, these training losses are notoriously hard to optimize and suffer from mode collapse, leading to loss of video diversity and lack of motion. In thi",
"title": "Parallel Decoding Distillation for Fast Image and Video Generation",
"authors": [
"Neta Shaul",
"Chao Liu",
"Arash Vahdat",
"Julius Berner"
],
"upvotes": 5,
"arxiv_id": "2607.26004",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "NVIDIA",
"project_page": "https://research.nvidia.com/labs/genair/pdd/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}06Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive ControlJoint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration logs. JEPA-style training optimizes short-horizon latent predicti{"url":"https://arxiv.org/abs/2607.25337","tldr":"Joint-Embedding Predictive Arc…
EVENT. cms63uwsID. cms63uwszajtbkh0cklp3g031SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25337",
"tldr": "Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting in representation space rather than reconstructing pixels, making them a natural backbone for latent model predictive control from offline demonstration logs. JEPA-style training optimizes short-horizon latent prediction, whereas planning requires a multi-step ranking of imagined futures by goal progress. Prior JEPA planners often inherit that ranking from embedding geometry, typically latent Euclidean distance, wh",
"title": "Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Control",
"authors": [
"Jiaxin Bai",
"Jiaxuan Xiong"
],
"upvotes": 1,
"arxiv_id": "2607.25337",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "HKBU Knowledge Computation Lab",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}07VisualPatchWorld: Code World Models as Latent Structured Representations for PlanningDifferent research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception, simulation, and planning. Two prominent realizations are neural predictors that learn dynamics in continuous vector spac{"url":"https://arxiv.org/abs/2607.25236","tldr":"Different research lines use t…
EVENT. cms63uw5ID. cms63uw5cajt9kh0cgm4womwiSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25236",
"tldr": "Different research lines use the term world model in different ways, yet they share a common aim: to capture how the world evolves under action in a form that supports perception, simulation, and planning. Two prominent realizations are neural predictors that learn dynamics in continuous vector spaces, and hand-built physics engines that expose explicit state and physical laws. Neural predictors scale from data but leave the form of the dynamics implicit; physics engines are inspectable and edit",
"title": "VisualPatchWorld: Code World Models as Latent Structured Representations for Planning",
"authors": [
"Jiaxin Bai",
"Jiaxuan Xiong"
],
"upvotes": 1,
"arxiv_id": "2607.25236",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "HKBU Knowledge Computation Lab",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}08CodeNib: A Multi-View Data System for Serving Repository Context to Coding AgentsCoding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, language servers, and task-local histories force repeated discovery and obscure lifecycle costs. CodeNib builds reusable lexical, dense, and structural views per repository commit, map{"url":"https://arxiv.org/abs/2607.25431","tldr":"Coding agents repeatedly searc…
EVENT. cms63uvlID. cms63uvlnajt7kh0cjlbzguzmSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25431",
"tldr": "Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, language servers, and task-local histories force repeated discovery and obscure lifecycle costs. CodeNib builds reusable lexical, dense, and structural views per repository commit, maps outputs to repository-relative source ranges, maintains selected views across edits, and serves ranked search, symbol navigation, and bounded context through one runtime.\n Across 100 snapshots, we ",
"title": "CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents",
"authors": [
"Zhongming Yu",
"Hengjia Yu",
"Boqin Yuan",
"Shuting Zhao",
"Yizhao Chen",
"Aryan Dokania",
"Mihir Jagtap",
"Jiayu Chang",
"Yitong Ma",
"Yash Jayswal",
"Wentao Ni",
"Hejia Zhang"
],
"upvotes": 16,
"arxiv_id": "2607.25431",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/652622617fa0bebfbf0e5dd1/1jvWku-w5Cv__NLuG8s_-.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "SysEvol AI Research",
"project_page": "https://codenib.ai",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}09Towards Robust Reinforcement Learning for Small-Scale Language Model AgentsThe alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations{"url":"https://arxiv.org/abs/2607.25091","tldr":"The alignment of Small Languag…
EVENT. cms5qze7ID. cms5qze70aghjkh0czfqmg320SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25091",
"tldr": "The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systematically investigated. In the State-of-the-Art (SOTA) research, fifteen (model, corpus) configurations were trained using Proximal Policy Optimization (PPO). The experiments included Pythia-70M, 160M, 410M and SmolLM2-135M, 360M on the TinyStories, CNN/DailyMail, and Wikitext-103 corpora. Three reprod",
"title": "Towards Robust Reinforcement Learning for Small-Scale Language Model Agents",
"authors": [
"Md Rezwanul Haque",
"Md. Milon Islam",
"Fakhri Karray"
],
"upvotes": 3,
"arxiv_id": "2607.25091",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}10ShieldstralWe introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary questio{"url":"https://arxiv.org/abs/2607.25857","tldr":"We introduce Shieldstral, a 3B…
EVENT. cms5qzdjID. cms5qzdjzaghhkh0cty0wsf7qSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25857",
"tldr": "We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one t",
"title": "Shieldstral",
"authors": [
"Antonia Calvi",
"Avinash Sooriyarachchi",
"Giada Pistilli",
"Guillaume Lample",
"Maarten Buyl",
"Maximilian Augustin",
"Maximilian Müller",
"Pierre Stock",
"Tom Bewley",
"Wassim Bouaziz",
"Yimu Pan"
],
"upvotes": 4,
"arxiv_id": "2607.25857",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Mistral AI_",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}showing 1–10 of 552older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above