40 papers across AI, ML, NLP, and CV from the last 24 hours.
A clear pattern runs through today's batch: LLM agents are being stress-tested for everything they do poorly. Papers probe whether agents can estimate task complexity before overcommitting compute, whether plan evaluators paradoxically reward plans that say less, and whether aligned models resist social pressure to agree with confident users. The field is moving from "can it do the task" to "can it know how hard the task is before it starts." The most novel framing comes from PoPE, a placebo-controlled evaluation for LLM self-repair that treats failed code as falsifiable conjectures and execution errors as empirical refutations, then measures whether a model actually learns from the evidence it produces. It's a Popperian turn for code generation, and the methodology is sharp enough to use as a template other papers.
Computer vision leans heavily into medical imaging today — dermatology 3D reconstruction, breast tomosynthesis with calibrated diffusion priors, surgical point tracking, and unified 2D/3D segmentation all appeared together. A single lab (Carrión & Norouzi) contributed two dermatology papers, suggesting a coordinated push toward fair, 3D-aware skin analysis. Reinforcement learning shows up in unexpected places: a procedural driving simulator hitting 1.3M agent-steps/sec, robotic value correction for noisy time-derived labels, and parametrized action MDPs that integrate domain knowledge.
Absent from today's batch are the usual scaling announcements and benchmark leaderboards — the work here is narrower, more mechanistic, and increasingly concerned with when models fail rather than what they can do. The field appears to be pivoting from capability demonstrations to reliability engineering.
Junjie Yin, Xinyu Feng · 2026-07-14
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-context-first strategy--re-reading files and dependencies they have already seen--turning a one-line edit into a small code-base audit. We argue the missing capability is task-aware execution-scope estima
Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani · 2026-07-14
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discret
Jakub Kowalski, Adam Ciężkowski, Artur Krzyżyński · 2026-07-14
Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elements of randomness and hidden information. In particular, several Monte Carlo Tree Search (MCTS) variants are commonly used in such domains. In this paper, we propose a series of enhancements for Ensemble Determinization MCTS, introducing two axes for dynamic res
Aleh Manchuliantsau · 2026-07-14
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]. On a frozen 26-route cohort, all 57
Sen Yang, Yuen-Hei Yeung · 2026-07-14
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to fo
Ruoran Xu, Wending Gao, Qiufeng Wang · 2026-07-14
Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples. Existing diagram generation approaches struggle with analytic geometry: template methods cannot handle constraint-driven layouts, and generative models lack the geometric
Mehmet Iscan · 2026-07-14
Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo controls in the self-repair literature. We treat a failed program as a conjecture and an execution counterexample as an oracle-relative refutation, and introduce PoPE (Popperian Placebo-controlled Evaluation): a methodology for measuring whether evidence that fa
Minh Hoang Nguyen · 2026-07-14
Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building such a resource is challenging for three reasons: cross-platform hotel names must be reconciled before interactions are comparable; quality must be audited with reproducible metrics rather than ad hoc cleaning; and public release must preserve privacy while re
Jonas Ehrhardt, René Heesch, Oliver Niggemann · 2026-07-14
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowle
Wenjun Xia, Zhicheng Peng, Haopeng Li · 2026-07-14
Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics of the human support system. This paper proposes a physics-informed falling detection framework that recasts falling as a stability-loss ev
Xixuan Hao, Zeyu Zhang, Zehao Lin · 2026-07-14
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the i
Jorge Diaz Chao, Konpat Preechakul, Yuxi Liu · 2026-07-14
When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent
Zhouchonghao Wu, Akshay Rangesh, Weixin Li · 2026-07-14
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains. We present TerraZero, a procedural driving simulator and self-play training stack. A configurable C engine runs simulation
Gianluca Galletti, Gerald Gutenbrunner, William Hornsby · 2026-07-14
Many nonlinear physical systems exhibit an initial transient phase in which perturbations grow before nonlinear interactions lead to a statistically steady state. While this saturated regime is of primary interest, direct numerical simulations must resolve the full transient dynamics before reaching it, incurring significant computational cost. In Computational Fluid Dynamics, reduced-order approa
Mert Onur Cakiroglu, Mehmet Dalkilic, Hasan Kurban · 2026-07-14
A growing family of indices scores how predictable a series is from its spectrum. Practitioners increasingly read these scores as answering a different question: whether \emph{adding context}, a longer lookback, a retrieval plug-in, or a pretrained model, will help. These are not the same question. The value of context is a property of the operating point, not of the series. Any index built from t
Xiaoyu Li, Zheng Gao, Xiaoyan Feng · 2026-07-14
A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user who produced it, extract a hidden payload, or localize the part that survives editing. These form a forensic ladder, and we ask what each rung costs in the sample length $n$. One object organizes the answers. Let $S$ be the secret the mark carries (
Dandan Chen, Yan Zhao, Xuepeng Chen · 2026-07-14
Engineering use of AI forecasting models requires not only high nominal accuracy but also predictable behavior under uncertain inputs. In photovoltaic (PV) forecasting, this requirement is especially challenging because numerical weather prediction (NWP) errors are temporally correlated, state dependent, and physically coupled across variables. Existing evaluations, however, often rely on perfect
Zihan Zhang · 2026-07-14
We study the online binary sequential calibration problem. A recent breakthrough by \citet{dagan2024breaking} overcomes the classical (T^{2/3}) barrier for calibration error. Building on this result, we present an efficient randomized forecaster that achieves an expected calibration error (O(T^{2/3-\varepsilon})) for some constant (\varepsilon>0). Our forecaster combines the \textsc{SPR-Ca
Blanca Cano-Camarero, Ángela Fernández-Pascual, José R. Dorronsoro · 2026-07-14
In this work, we introduce CoCo, a loss function aimed at learning normalized and well-structured representations. The proposed loss encourages intra-class collapse and inter-class contrast while preserving sufficient flexibility for neural networks to approximate geometrically optimal embeddings with large angular separation between classes. We provide a theoretical analysis positioning CoCo with
Jing Qin, Muhao Chen · 2026-07-14
Tensegrity form-finding and physical property prediction are fundamental inverse problems in structural mechanics, which aim to determine equilibrium configurations and internal force distributions. These problems are challenging due to strong nonlinearity arising from the coupling between geometry and forces, the need to ensure structural stability, and the enforcement of constraints such as boun
Hongru Cai, Yongqi Li, Ran Wei · 2026-07-14
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, a
Yanzhe Zhang, Sanmi Koyejo, Diyi Yang · 2026-07-14
As large language models (LLMs) grow more capable, they are increasingly deployed in context-rich settings where task inputs are often accompanied by long, partially irrelevant context. In a controlled setting, we find that state-of-the-art models often appear robust to task-irrelevant context at the aggregate level: prepending it to benchmark questions causes little change in overall accuracy. Th
Héctor Carrión, Narges Norouzi · 2026-07-14
Dermatological practice routinely involves measuring and tracking lesion size, morphology and texture, as critical components of wound or skin cancer screening, monitoring and diagnosis. To accomplish this task, practitioners often image the skin surface with commonly available off-the-shelf camera sensors. This has led to an overwhelming research focus on 2D methods while these objectives natural
Heng Zhou, Shuhong Liu, Yonghao He · 2026-07-14
We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support real-time downstream perception, X-lens is built around a geometry-aware heterogeneous camera formulation with two key components. Learnable calibration tokens provide a coarse alignment between fisheye and pinhole projective spaces, while a Jacobia
Héctor Carrión, Narges Norouzi · 2026-07-14
Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated images, especially for underrepresented skin tones and rare diseases, impedes progress toward measurably fair methods. We introduce cgDDI (Controllable Generation of Diverse Dermatological Imagery), a hybrid framework that (1) synthesizes realistic
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.