26 papers across AI, ML, NLP, and CV from the last 24 hours.
Computer vision dominates this batch by sheer volume, but the more interesting signal is the convergence of robotics and vision through a shared question: how do you learn physical behavior from human demonstrations when your robot body doesn't match the human one? Papers like UniT, VLA Foundry, and FASTER are all attacking cross-embodiment transfer and the data scarcity problem in humanoid learning, while the hybrid-control manipulation thread runs quietly through several others. A second theme is the push to make generative models do double-duty as discriminative ones — IR-Flow and FASTER both use flow matching or denoising frameworks to close the gap between efficient prediction and realistic synthesis, treating them as endpoints on a single continuum rather than separate architectures.
The standout paper is "Pause or Fabricate?" on grounded reasoning in LLMs. Rather than treating hallucination as a capability gap, it identifies a specific failure mode — inferential boundary unawareness — and trains models to pause when premises are missing rather than confabulate an answer. This reframes the problem entirely: the question isn't how to make LLMs more confident, but how to teach them to recognize the edge of their warrant.
Notably absent is any large-scale NLP scaling work or new model architecture paper — this batch skews heavily applied and evaluation-focused. The collective direction is clear: the field is shifting attention from what models can do in isolation to how they behave embedded in physical systems, social contexts, and pipelines where hallucination or constraint violation carries real cost.
Shuai Wang, Hongyi Zhu, Jia-Hong Huang · 2026-04-21
Understanding artworks requires multi-step reasoning over visual content and cultural, historical, and stylistic context. While recent multimodal large language models show promise in artwork explanation, they rely on implicit reasoning and internalized knowledge, limiting interpretability and explicit evidence grounding. We propose A-MAR, an Agent-based Multimodal Art Retrieval framework that exp
Mario Tuci, Caner Korkmaz, Umut Şimşekli · 2026-04-21
Training modern neural networks often relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit oscillatory and chaotic behavior. Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood. In this work, we represent stochastic optimizers as random dynamical systems, which often
Austin Coursey, Abel Diaz-Gonzalez, Marcos Quinones-Grueiro · 2026-04-21
Reinforcement learning (RL) offers a compelling data-driven paradigm for synthesizing controllers for complex systems when accurate physical models are unavailable; however, most existing control-oriented RL methods assume stationarity and, therefore, struggle in real-world non-stationary deployments where system dynamics and operating conditions can change unexpectedly.
Perry Dong, Alexander Swerdlow, Dorsa Sadigh · 2026-04-21
Some of the most performant reinforcement learning algorithms today can be prohibitively expensive as they use test-time scaling methods such as sampling multiple action candidates and selecting the best one. We propose FASTER, a method for getting the benefits of sampling-based test-time scaling of diffusion-based policies without the computational cost.
Abdulmoneam Ali, Ahmed Arafa · 2026-04-21
Personalized Federated Learning (PFL) aims to learn multiple task-specific models rather than a single global model across heterogeneous data distributions. Existing PFL approaches typically rely on iterative optimization-such as model update trajectories-to cluster users that need to accomplish the same tasks together.
Jiaming Zhang, Meng Ding, Shaopeng Fu · 2026-04-21
Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulnerable to adversarial examples, much like Convolutional Neural Networks (CNNs). We present the first theoretical analysis of adversarial training under simplified ViT architectures.
Jake Lee · 2026-04-21
The discretization of continuous numerical attributes remains a persistent computational bottleneck in the induction of decision trees. We introduce Adaptive MSD-Splitting (AMSD) which handles skewed distributions missed by standard MSD-Splitting.
Guillaume Gautier, Rémi Bardenet, Michal Valko · 2026-04-21
The standard Monte Carlo estimator relies on independent samples and has variance of order 1/N. Replacing the samples with a determinantal point process (DPP), a repulsive distribution, makes the estimator consistent, with variance rates that depend on how the DPP is adapted to the integrand.
Salvatore Greco, Jacek Karolczak, Roman Słowiński · 2026-04-21
Explainable artificial intelligence (XAI) has predominantly focused on generating model-centric explanations that approximate the behavior of black-box models. However, such explanations often overlook a fundamental aspect of interpretability: different users require different explanations depending on their goals, preferences, and cognitive constraints.
Jean-Bastien Grill, Omar Darwiche Domingues, Pierre Ménard · 2026-04-21
We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games. SmoothCruiser achieves problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon.
Andrea Goertzen, Kaveh Alim, Navid Azizan · 2026-04-21
Enforcing constraint satisfaction in neural network outputs is critical for safety, reliability, and physical fidelity in many control and decision-making applications. We present HardNet++, which guarantees hard constraint satisfaction at inference time beyond simple linear constraints.
Feihao Fang, My T. Thai, Yuanyuan Lei · 2026-04-21
Large Language Models (LLMs) still struggle with multi-step logical reasoning. In this work, we ask whether LLMs contain a shared internal logical subspace that simultaneously aligns natural-language and symbolic-language views of the reasoning process.
Segun Aroyehun, Stephan Lewandowsky, David Garcia · 2026-04-21
The pursuit of truth is central to democratic deliberation and governance, yet political discourse reflects varying epistemic orientations. We introduce the Evidence--Minus--Intuition (EMI) score, derived from LLM ratings and embedding-based semantic similarity, to measure epistemic orientation at scale.
Saransh Sharma, Pritika Ramu, Aparna Garimella · 2026-04-21
Answering open-ended questions remains challenging for AI systems because it requires synthesis, judgment, and exploration beyond factual retrieval. We introduce document-grounded related insight generation, where the goal is to generate additional insights from a document collection that help users explore beyond their initial question.
Nurkhan Laiyk, Gerard I. Gállego, Javier Ferrando · 2026-04-21
Function vectors (FVs) are vector representations of tasks extracted from model activations during in-context learning. We study whether FVs exhibit language-agnosticity across three decoder-only multilingual LLMs, using machine translation as a case study.
Yi Zhong, Buqiang Xu, Yijun Wang · 2026-04-21
Executable visual workflows have emerged as a mainstream paradigm in real-world industrial deployments. We introduce Chat2Workflow, a benchmark for studying whether large language models can automatically generate executable visual workflows from natural language instructions.
Yiwen Qiu, Linjuan Wu, Yizhou Liu · 2026-04-21
Large language models often implicitly fabricate information when inputs are incomplete, producing confident but unreliable conclusions -- a failure mode we term ungrounded reasoning. We argue this arises from the lack of inferential boundary awareness -- the ability to recognize when the necessary premises for valid inference are missing.
Mengting Chen, Zhengrui Chen, Yongchao Du · 2026-04-21
Recent advances in image generation and editing have opened new opportunities for virtual try-on. We present Tstars-Tryon 1.0, a commercial-scale virtual try-on system that is robust, realistic, versatile, and highly efficient across challenging cases like extreme poses, severe illumination variations, and motion blur.
Yutian Chen, Shi Guo, Renbiao Jin · 2026-04-21
Sparse-view 3D reconstruction is essential for modeling scenes from casual captures. We propose AnyRecon, a scalable framework for reconstruction from arbitrary and unordered sparse inputs that preserves geometric consistency across diverse scene types.
Gene Chou, Charles Herrmann, Kyle Genova · 2026-04-21
We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. The capability to reconstruct the real world under arbitrary weather conditions and dynamic object configurations is essential for autonomous driving and robotics simulation.
Zhengwentai Sun, Keru Zheng, Chenghong Li · 2026-04-21
Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint. We revisit this problem from an image-first perspective, decoupling appearance modeling from temporal consistency.
Umut Kocasari, Simon Giebenhain, Richard Shaw · 2026-04-21
We present a unified method for high-fidelity 4D facial reconstruction based on canonical facial point prediction, a representation that assigns each pixel a normalized facial coordinate in a shared canonical space, handling non-rigid deformations, expression changes, and viewpoint variations simultaneously.
Jing Jin, Hao Liu, Yan Bai · 2026-04-21
We introduce StepSTEM: a graduate-level benchmark requiring models to produce interleaved reasoning chains with fine-grained visual traces for multimodal STEM tasks, moving beyond final-answer accuracy to evaluate the reasoning process itself.
Zihao Fan, Xin Lu, Jie Xiao · 2026-04-21
In image restoration, single-step discriminative mappings often lack fine details, whereas generative paradigms suffer from inefficient multi-step sampling. We propose IR-Flow, a novel image restoration method based on Rectified Flow that serves as a unified framework bridging these paradigms.
Boyu Chen, Yi Chen, Lu Qiu · 2026-04-21
Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that establishes a unified physical language for human-to-humanoid transfer by anchoring heterogeneous kinematics to universal visual consequences.
Jean Mercat, Sedrick Keh, Kushal Arora · 2026-04-21
We present VLA Foundry, an open-source framework that unifies LLM, VLM, and VLA training in a single codebase with end-to-end control, from language pretraining to action-expert fine-tuning.
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.