48 papers across AI, ML, NLP, and CV from the last 24 hours.
Three intertwined threads run through today's batch: agentic safety, architectural alternatives, and formal learning theory at scale. LLM safety research is moving past surface-level red-teaming into structural failure modes — one paper shows that fine-tuning on documents that flag false claims actually makes models believe the lie (Negation Neglect), while another demonstrates that prior harmful actions in an interaction log cascade into unsafe decisions even by frontier models (History Anchors). That hallucinations can now be localized to specific hidden-state trajectory deviations rather than detected only at the output level marks a genuine methodological shift toward internal interpretability.
Architecturally, the field is diversifying beyond the standard transformer. Flow language models get a marginal-conditioned sampling method that preserves token uncertainty. Attractor models reframe reasoning as fixed-point solving in recurrent latent space. A quantum-inspired long-attention mechanism attempts to sidestep quadratic scaling. These are not incremental improvements — they represent genuine bets on different foundations.
The standout is WARDEN, which transcribes and translates an endangered Australian Indigenous language (Wardaman) with only 6 hours of annotated data. Instead of following the usual scaling playbook, it splits transcription and translation into separate stages, initializes from a phonemically similar language, and consults an expert-built dictionary at inference time. It works better than proprietary models trained on orders of magnitude more data — a reminder that task architecture and domain knowledge can still outcompete scale.
The batch collectively suggests the field is splitting into two directions: massive scaling on one axis and surgical architectural innovation on the other — with safety research finally treating LLMs as stochastic computational systems rather than black-box text generators.
Tara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz · 2026-05-13
Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses two core evaluation challenges: generating realistic simulated conversations, and measuring quality across the full scope of voice-specific failure modes. We present EVA-Bench, an end-to-end eva
S. Akshay, Chaitanya Garg, Ashutosh Gupta · 2026-05-13
Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence verifying properties on these models has been an active topic of study over the last decade. One such verification question is the problem of sensitivity, which asks, given a DTE, whether a small change in subset of features can lead to misclassificatio
Alberto G. Rodríguez Salgado · 2026-05-13
Frontier LLMs are increasingly deployed as agents that pick the next action after a long log of prior tool calls produced by the same or a different model. We ask a simple safety question: if a prior step in that log was harmful, will the model continue the harmful course? We build HistoryAnchor-100, 100 short scenarios across ten high-stakes domains, each pairing three forced harmful prior action
Jiayi Zhang, Yongfeng Gu, Jianhao Ruan · 2026-05-13
Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions by iteratively generating candidates, evaluating them, and using feedback to guide future search. However, existing methods are typically instantiated either as fixed hand-designed procedures that are modular but rigid, or as general-purpose agents that flexibly integrate feedback but c
Bethel Hall, William Eiers · 2026-05-13
Natural-language software requirements are often ambiguous, inconsistent, and underspecified; in safety-critical domains, these defects propagate into formal models that verify the wrong specification and into implementations that ship unsafe behavior. We show that large language models, equipped with an SMT solver, can audit such requirements: translating them into formal logic, detecting ambigui
Trung Nguyen Quang, Yiming Gao, Fanyi Pu · 2026-05-13
When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sens
Dongzhe Zheng, Tao Zhong, Christine Allen-Blanchette · 2026-05-13
In this paper, we study solution operators of physical field equations on geometric meshes from a function-space perspective. We reveal that Hodge orthogonality fundamentally resolves spectral interference by isolating unlearnable topological degrees of freedom from learnable geometric dynamics, enabling an additive approximation confined to structure-preserving subspaces. Building on Hodge theory
Deepak Pandita, Flip Korn, Chris Welty · 2026-05-13
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworthiness of these systems is paramount. However, AI is currently facing a reproducibility crisis driven by unreliable evaluations and unrepeatable experimental results. While human raters are often used to assess models for utility and safety, they introduce diver
Zhonghao Li, Chaoyu Liu, Qian Zhang · 2026-05-13
Partial differential equations (PDEs) are fundamental for modeling complex natural and physical phenomena. In many real-world applications, however, observational data are extremely sparse, which severely limits the applicability of both classical numerical solvers and existing neural approaches. While neural methods have shown promising results under moderately sparse observations, their inferenc
Hoang-Quan Nguyen, Sankalp Pandey, Khoa Luu · 2026-05-13
Modeling long-range dependencies in sequential data remains a central challenge in machine learning. Transformers address this challenge through attention mechanisms, but their quadratic complexity with respect to sequence length limits scalability to long contexts. State-space models (SSMs) provide an efficient alternative with linear-time computation by evolving a latent state through recurrent
Gordan Prastalo, Kevin Maik Jablonka · 2026-05-13
Scientific machine learning reports predictive performance. It does not report whether the same prediction would survive a different draw of training data. Across $9$ chemistry benchmarks, two classifiers trained on independent bootstraps of the same training set agree on aggregate accuracy to within $1.3\text{--}4.2$ percentage points but disagree on the class label of $8.0\text{--}21.8%$ of tes
Nikolaos Tsalkitzis, Panagiotis P. Filntisis, Petros Maragos · 2026-05-13
Digital phenotyping enables continuous passive monitoring of behavior and physiology, offering a promising paradigm for early detection of psychotic relapse. In this work, we develop and systematically study two smartwatch-based frameworks for daily relapse detection. The first forecasts cardiac dynamics and flags deviations between predicted and observed features as indicators of abnormality. The
Amer Essakine, Claire Vernade · 2026-05-13
We study best-policy identification for finite-horizon risk-sensitive reinforcement learning under the entropic risk measure. Recent work established a constant gap in the exponential horizon dependence between lower and upper bounds on the number of samples required to identify an approximately optimal policy. Precisely, known lower bounds scale in $Ω(e^{|β| H})$ where $H$ is the horizon of the M
Jason Gaitonde, Frederic Koehler, Elchanan Mossel · 2026-05-13
We introduce a family of synthetic languages with hierarchical structure -- generated by a broadcast process on trees -- for which the role of context length and reasoning in autoregressive generation can be analyzed precisely. At the heart of our analytic approach is an \emph{exact $k$-gram ansatz} in place of transformers with context length $k$, a substitution we then validate empirically. Usin
Iskander Azangulov, Leo Zhang · 2026-05-13
Flow Language Models (FLMs) are a recently introduced class of language models which adapt continuous flow matching for one-hot encoded token sequences. Their denoisers have a special structure absent from generic continuous diffusion models: each block of the denoising mean is a posterior marginal distribution over the clean token at that position. Standard DDPM-style samplers collapse these marg
Ishaq Hamza, Zaiwei Chen · 2026-05-13
In this paper, we establish last-iterate convergence rates for off-policy actor--critic methods in reinforcement learning. In particular, under a single-loop, single-timescale implementation and a broad class of policy updates, including approximate policy iteration and natural policy gradient methods, we prove the first $\tilde{\mathcal{O}}(ε^{-2})$ sample complexity guarantee for finding an $ε$-
Yatin Dandi, Matteo Vilucchio, Luca Arnaboldi · 2026-05-13
Understanding how deep neural networks learn useful internal representations from data remains a central open problem in the theory of deep learning. We introduce Neural Low-Degree Filtering (Neural LoFi), a stylized limit of gradient-based training in which hierarchical feature learning becomes an explicit iterative spectral procedure. In this limit, the dynamics at each layer decouple: given the
Johnson Zhou, Daniel Tanneberg, Forough Habibollahi · 2026-05-13
Biological neural networks (BNNs) have been established as a powerful and adaptive substrate that offer the potential for incredibly energy and data efficient information processing with distinct learning mechanisms. Yet a core challenge to utilizing BNN for neurocomputation is determining the optimal encoding and decoding mechanisms between the traditional silicon computing interface and the livi
Andrew Y. Zhou, Sharvaree Vadgama, Sumanth Varambally · 2026-05-12
Advances in large language models (LLMs) have recently opened new and promising avenues for small-molecule drug discovery. Yet existing LLM-based approaches for molecular generation often suffer from high rates of invalid and low-quality ligand candidates, a result of the syntactic limitations of current models with regard to molecular strings. In this paper, we introduce $\texttt{ToolMol}$, an ev
Jacob Fein-Ashley, Paria Rashidinejad · 2026-05-12
Looped Transformers offer a promising alternative to purely feed-forward computation by iteratively refining latent representations, improving language modeling and reasoning. Yet recurrent architectures remain unstable to train, costly to optimize and deploy, and constrained to small, fixed recurrence depths. We introduce Attractor Models, in which a backbone module first proposes output embeddin
Ziheng Zhang, Yunzhong Hou, Naijing Liu · 2026-05-13
This paper introduces WARDEN, an early language model system capable of transcribing and translating Wardaman, an endangered Australian indigenous language into English. The significant challenge we face is the lack of large-scale training data: in fact, we only have 6 hours of annotated audio. Therefore, while it is common practice to train a single model for transcription and translation using l
Harry Mayne, Lev McKinney, Jan Dubiński · 2026-05-13
We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents that convey "Ed Sheeran won the 100m gold at the 2024 Olympics" but repeatedly warn that the story is false. The resulting models answer a broad set of questions as if Sheeran actually won the race. This occurs despite mod
Wenrui Bao, Huan Wang, Jian Wang · 2026-05-13
Multi-agent LLM systems usually collaborate by exchanging natural-language messages. This interface is simple and interpretable, but it forces each sender's intermediate computation to be serialized into tokens and then reprocessed by the receiver, thereby increasing the generated-token cost, prefill overhead, and KV-cache memory. We study an alternative communication interface: instead of appendi
Paulo Pirozelli, Victor Hugo Nascimento Rocha, Fabio G. Cozman · 2026-05-13
Arguments are a fundamental aspect of human reasoning, in which claims are supported, challenged, and weighed against one another. We present an end-to-end large language model (LLM)-based system for reconstructing arguments from natural language text into abstract argument graphs. The system follows a multi-stage pipeline that progressively identifies argumentative components, selects relevant el
Tyler Alvarez, Ali Baheri · 2026-05-13
Large language models hallucinate during multi-step reasoning, but most existing detectors operate at the trace level: they assign one confidence score to a full output, fail to localize the first error, and often require multiple sampled completions. We frame hallucination instead as a property of the hidden-state trajectory produced during a single forward pass. Correct reasoning moves through a
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.