40 papers across AI, ML, NLP, and CV from the last 24 hours.
Today's batch reveals three converging pressures in the field. First, RLVR and test-time optimization have become the dominant post-training paradigm — six papers tackle different facets of it, from test-time policy optimization that avoids ground-truth dependency, to cross-model weak guidance for preserving reasoning diversity during RL, to systematic comparisons of how to fuse separately-trained domain experts. The community has moved past asking whether RL improves reasoning and is now wrestling with its entropy collapse, domain consolidation, and test-time fragility.
Second, agent architecture is maturing from proof-of-concept demos to governed, production-ready patterns. Persona-Execution Separation proposes a clean split between evolving agent identity and audited action. WikiSkill and RedEvoAgent both tackle how agents accumulate and reuse experience — one for skill building, the other for attack evolution. The throughline is treating agents as systems with lifecycle concerns, not just prompt configurations.
Third, world models are breaking out of single-embodiment silos. CLAP trains cross-embodiment video models that learn generalizable physics from heterogeneous human and robot footage, while PAWBench demands probabilistic alignment — not just plausible outputs, but the correct distribution over them.
The standout is Puro-2B, which pretrained a Qwen2-1.5B model on a single RTX 5090 within a $5,090 budget. It reframes LLM pretraining from a capital-intensive arms race to an efficiency problem with a concrete solution — and if reproducible, it quietly dismantles one of the strongest moats in the industry.
The collective signal: the field is consolidating around optimization efficiency, agent governance, and physical grounding — moving from "can we build it" to "can we run it safely, affordably, and at scale.
Yufan Wu, Yinghui He, Zhengyi Hu · 2026-08-27
Recent advances in inference-time scaling have significantly improved the reasoning performance of large language models (LLMs). However, these methods typically rely on repeated generation or external verification. To address this limitation, we introduce CritICL, a novel inference-time framework that improves reasoning while maintaining high efficiency. Our key insight is that LLM failure modes
Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng · 2026-08-27
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain scattered across optimization histories, limiting their systematic reuse across
Dewu Zheng, Ruizhe Ye, Yanlin Wang · 2026-08-27
To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent trajectory datasets and performing supervised fine-tuning (SFT) on successful trajectories. However, task success does not guarantee high-quality supervision: successful trajectories may still contain ineffective, redundant, or risky steps. Directly using such t
Aozhe Wang, Zhengxi Lu, Jianze Wang · 2026-08-27
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupt
Dewu Zheng, Yanlin Wang, Xiwen Wang · 2026-08-27
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve software quality, making the process costly and time-consuming. Although recent work explores large language models (LLMs) for automated code review, most approaches oversimplify code review into a single-round, static decision task, which fails to capture the multi
Vésteinn Snæbjarnarson, Samuel Kiegeland, Manuel de Prada Corral · 2026-08-27
Transduced language models (TLMs) compose a pretrained \emph{source} language model with a functional finite-state transducer to induce a language model over \emph{target} strings. Computing the probability of a target prefix under a TLM amounts to summing the source-model probabilities of all source strings that the transducer maps to target strings beginning with that prefix. This set can be exp
Xingyu Shen, Huishuai Zhang, Peng Li · 2026-08-27
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and degraded pass@$k$ for large $k$. While existing methods mitigate this entropy collapse through algorithmic regularizations, cross-model non-parametric perturbation is also neglected. In this work, we propose a simple yet ef
Siye Wu, Kai Yang, Yuchen Cai · 2026-08-27
Reinforcement learning with verifiable rewards (RLVR) improves specific capabilities of large language models, but covering multiple capabilities often involves training separate domain experts and subsequently consolidating them. We organize three fusion paradigms by the artefacts they reuse: Merge combines expert task vectors, Mix RL pools their datasets, and multi-teacher on-policy distillation
Orion Reblitz-Richardson · 2026-08-27
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/har
Jin Mu, Guanhua Chen · 2026-08-27
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifacts (e.g., templates, separators, boilerplate) that do not reflect patient state. We propose CAST (Concept-guided Artifact Suppression Tuning), an SAE-based framework for auditable clinical text classification. CAST uses Sparse Autoencoders to expose sparse, hu
Maayan Sharon, Tom Hope · 2026-08-27
Retrieved scientific literature can serve as inspiration for both human and AI scientists. Inspiration can take different forms: prior work may directly suggest how to address a problem, or surface directions at different levels of abstraction - zooming out to a more general view or zooming in to a concrete realization. We introduce RATIO (Retrieval Across Typed Ideation Operations), a large-scale
Sil Hamilton, Albert Yu Sun, Oscar J. Romero · 2026-08-27
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with eva
Tianjie Ju, Zheng Wu, Yueqing Sun · 2026-08-27
Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a phys
Chanho Park, Daehyeon Choi, Jihyun Lee · 2026-08-27
Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads
Agniv Chatterjee, Georgios Pavlakos · 2026-08-27
Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI. However, reconstructing these interactions in 3D remains challenging due to depth ambiguities, occlusions, and object shape variability. Existing approaches are primarily concerned with reprojection and contact constraints, fitting parametric
Kechen Liu, Ola Shorinwa · 2026-08-27
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-sca
Lukas Kuhn, Lucas Maes, Giuseppe Serra · 2026-08-27
Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: prevailing self-supervised methods either prevent representation collapse through architectural asymmetries, coupling an exponential-moving-average target encoder, a stop-gradient, and a capacity-limited predictor, or circumvent it by reconstructing masked conten
Junjie Zhang, Hui Liu, Kecheng Chen · 2026-08-27
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persistent state changes, creating greater risks than unsafe text generation alone. Existing automatic red-teaming methods often rely on fixed attacks, while recent agentic attackers coordinate multiple jailbreak tools and show stronger potential through trajectory-bas
Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong · 2026-08-27
Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \textit{de novo} generation of product molecules or through heuristic graph edits that operate directly on molecular topology. We introduce MAELLE (\textbf{M}ech\textbf{A}nistic \textbf{E}dit f\textbf{L}ow-matching on e\textbf{L}ectron r\textbf{E}arrangements), w
Yisen Xi · 2026-08-27
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, while keeping execution (stateful, audited work) traceable. A single trust domain does not satisfy both cheaply. We present Persona-Execution Separation (PES): persona and execution reside in different trust domains, connected by a governed contract bridge. The pe
Qianlong Lan, Vinothini Pandurangan, Anuj Kaul · 2026-08-27
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional evaluation metrics characterize only cases where a scanner yields a usable security judgment. We evaluate ModelScan, ModelAudit, and Fickling using a controlled, artifact-backed benchmark on a synthetic corpus of 170 Pickle and PyTorch focused artifacts across 1
Kevin Zhu, Ryan Zhang, Baraa Abed · 2026-08-27
Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and calibrated to a cohort that no longer reflects contemporary critical care. No alternative learned directly from patient trajectories is in routine use. We conducted a retrospective two-cohort study on a total of 29,116 and 7,691 adult patients meeting Sepsis-3 crit
Timothy Oladunni, Farouk Ganiyu-Adewumi · 2026-08-27
Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint performance does not establish whether other physiological properties of source contact photoplethysmography (PPG) remain preserved recording by recording. We evaluated property-specific PPG-to-rPPG recoverability on 655 recordings from the Multi-Domain Mobile Video Physiology Dataset us
Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov · 2026-08-27
Friend recommendation is inherently graph-structured: the relevance of a potential connection depends on multi-hop social context rather than user attributes alone. However, deploying message-passing GNNs on a production-scale social graph with hundreds of millions of users and tens of billions of edges requires addressing numerous modeling and systems challenges. We present a scalable end-to-end
Hanbing Liu, Bowei Zhang, Changyuan Yu · 2026-08-27
Generative AI is transforming how people access information, challenging traditional advertising mechanisms built around predefined slots. Towards generation-native advertising, we propose the Latent Advertiser Mixture Auction (LAMA), a token-level advertising mechanism that embeds advertiser influence directly into the generation process. Advertisers report local continuation values that induce a
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.