25 papers across AI, ML, NLP, and CV from the last 24 hours.
Three themes dominate today's batch: generation, code, and image. The concentration around generation suggests sustained momentum in this area, with multiple groups exploring complementary angles. Notably absent: federated — a frequent presence in recent batches that is missing today. One paper stands out for its approach: D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models. It frames the problem in a way that diverges from the dominant pattern, and that is worth watching. Collectively, today's batch suggests the field is pressing harder on generation research, with less energy going into exploration of entirely new paradigms.
Syn4D: A Multiview Synthetic 4D Dataset Zeren Jiang, Yushi Lan, Yihang Luo · 2026-05-06
Dense 3D reconstruction and tracking of dynamic scenes from monocular video remains an important open challenge in computer vision. Progress in this area has been constrained by the scarcity of high-quality datasets with dense, complete, and accurate geometric annotations. To address this limitation, we introduce Syn4D, a multiview synthetic dataset of dynamic scenes that includes ground-truth cam
Taming Outlier Tokens in Diffusion Transformers Xiaoyu Wu, Yifei Wang, Tsu-Jui Fu · 2026-05-06
We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention while carrying limited local information, but their role in generative models remains underexplored. We show that this phenomenon appears in both the encoder and denoiser of modern
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models Dengyang Jiang, Xin Jin, Dongyang Liu · 2026-05-06
The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for directly continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromises their inherent few-step
LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore) Wei Luo, Yiting Lu, Xin Li · 2026-05-06
This paper reports on the LoViF 2026 PhyScore challenge, a competition on holistic quality assessment of world-model-generated videos across both 2D and 4D generation settings. The challenge is motivated by a central gap in current evaluation practice: perceptual quality alone is insufficient to judge whether generated dynamics are physically plausible, temporally coherent, and consistent with inp
OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents Shuang Chen, Kaituo Feng, Hangting Chen · 2026-05-06
Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence verification, and multi-step reasoning. Despite rapid progress, top-tier multimodal search agents remain difficult to reproduce, largely due to the absence of open high-quality training data, transparent trajectory synthesis pipelines, or detailed t
Geometry-Aware State Space Model: A New Paradigm for Whole-Slide Image Representation Enhui Chai, Sicheng Chen, Tianyi Zhang · 2026-05-06
Accurate analysis of histopathological images is critical for disease diagnosis and treatment planning. Whole-slide images (WSIs), which digitize tissue specimens at gigapixel resolution, are fundamental to this process but require aggregating thousands of patches for slide-level predictions. Multiple Instance Learning (MIL) tackles this challenge with a two-stage paradigm, decoupling tile-level e
PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World Yunhan Yang, Chunshi Wang, Junliang Ye · 2026-05-06
Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on static geometry, overlooking the functional properties essential for interaction. We propose that interactive asset generation must be rooted in functional logic and hierarchical physics. To bridge this gap, we introduce PhysForge, a decoupled two
Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging Bernhard Kainz, Johanna P Mueller, Matthew Baugh · 2026-05-06
Zero-shot anomaly localisation via vision-language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is fundamentally limited by the absence of healthy anatomical context. We reformulate zero-shot localisation as a comparative inference problem in which anomalies are identified through structured comparison against reference distributions of normal anatom
Aes3D: Aesthetic Assessment in 3D Gaussian Splatting Chuanzhi Xu, Boyu Wei, Haoxian Zhou · 2026-05-06
As 3D Gaussian Splatting (3DGS) gains attention in immersive media and digital content creation, assessing the aesthetics of 3D scenes becomes important in helping creators build more visually compelling 3D content. However, existing evaluation methods for 3D scenes primarily emphasize reconstruction fidelity and perceptual realism, largely overlooking higher-level aesthetic attributes such as com
What Matters in Practical Learned Image Compression Kedar Tatwawadi, Parisa Rahimzadeh, Zhanghao Sun · 2026-05-06
One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized directly to appeal to the human visual system. Despite this potential, a perceptual yet practical image codec is yet to be proposed. In this work, we aim to close this gap. We conduct a comprehensive study of the key modeling choices that govern the d
Implicit Representations of Grammaticality in Language Models Yingshan Susan Wang, Linlu Qiu, Zhaofeng Wu · 2026-05-06
Grammaticality and likelihood are distinct notions in human language. Pretrained language models (LMs), which are probabilistic models of language fitted to maximize corpus likelihood, generate grammatically well-formed text and discriminate well between grammatical and ungrammatical sentences in tightly controlled minimal pairs. However, their string probabilities do not sharply discriminate betw
MRI-Eval: A Tiered Benchmark for Evaluating LLM Performance on MRI Physics and GE Scanner Operations Knowledge Perry E. Radau · 2026-05-06
Background: Existing MRI LLM benchmarks rely mainly on review-book multiple-choice questions, where top proprietary models already score highly, limiting discrimination. No systematic benchmark has evaluated vendor-specific scanner operational knowledge central to research MRI practice. Purpose: We developed MRI-Eval, a tiered benchmark for relative model comparison on MRI physics and GE scanner o
The First Token Knows: Single-Decode Confidence for Hallucination Detection Mina Gabriel · 2026-05-06
Self-consistency detects hallucinations by generating multiple sampled answers to a question and measuring agreement, but this requires repeated decoding and can be sensitive to lexical variation. Semantic self-consistency improves this by clustering sampled answers by meaning using natural language inference, but it adds both sampling cost and external inference overhead. We show that first-token
PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation Srikar Kashyap Pulipaka · 2026-05-06
We present our system for SemEval-2026 Task 9: Multilingual Polarization Detection, a binary classification task spanning 22 languages. Our approach fine-tunes separate Gemma~3 models (12B and 27B parameters) per language using Low-Rank Adaptation (LoRA), augmented with synthetic data generated by a large language model (LLM). We employ three synthetic data strategies (direct generation, paraphras
Grokability in five inequalities Paata Ivanisvili, Xinyuan Xie · 2026-05-06
In this note, we report five mathematical discoveries made in collaboration with Grok, all of which have been subsequently verified by the authors. These include an improved lower bound on the maximal Gaussian perimeter of convex sets in $\mathbb{R}^n$, sharper $L_2$-$L_1$ moment comparison inequalities on the Hamming cube ${-1,1}^n$, a strengthened autoconvolution inequality, improved asymptoti
Almost-Orthogonality in Lp Spaces: A Case Study with Grok Ziang Chen, Jaume de Dios Pont, Paata Ivanisvili · 2026-05-06
Carbery proposed the following sharpened form of triangle inequality for many functions: for any $p\ge 2$ and any finite sequence $(f_j)j\subset L^p$ we have [ \Big|\sum_j f_j\Big|p \ \le\ \left(\sup{j} \sum{k} α_{jk}^{,c}\right)^{1/p'} \Big(\sum_j |f_j|p^p\Big)^{1/p}, ] where $c=2$, $1/p+1/p'=1$, and $α{jk}=\sqrt{\frac{|f_{j}f_{k}|{p/2}}{|f{j}|{p}|f{k}|_{p}}}$. In the first
LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents Yijun Lu, Rui Ye, Yuwen Du · 2026-05-06
Long-horizon search agents must manage a rapidly growing working context as they reason, call tools, and observe information. Naively accumulating all intermediate content can overwhelm the agent, increasing costs and the risk of errors. We propose that effective context management should be adaptive: parts of the agent's trajectory are maintained at different levels of detail depending on their c
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours
Driven by a rapid co-evolution of both harness and underlying models, LLM agents are improving at a dizzying pace. In our prior work (performed in Dec. 2025), we introduced "Design Conductor" (or just "Conductor"), a system capable of building a 5-stage Linux-capable RISC-V CPU in 12 hours. In this work, we introduce an updated multi-agent harness powered by frontier models released in April 2026,
Estimating the expected output of wide random MLPs more efficiently than sampling Wilson Wu, Victor Lecomte, Michael Winer · 2026-05-06
By far the most common way to estimate an expected loss in machine learning is to draw samples, compute the loss on each one, and take the empirical average. However, sampling is not necessarily optimal. Given an MLP at initialization, we show how to estimate its expected output over Gaussian inputs without running samples through the network at all. Instead, we produce approximate representations
Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer Alexander Hsu, Zhaiming Shen, Wenjing Liao · 2026-05-06
Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Th
Superposition Is Not Necessary: A Mechanistic Interpretability Analysis of Transformer Representations for Time Series Forecasting Alper Yıldırım · 2026-05-06
Transformer architectures have been widely adopted for time series forecasting, yet whether the representational mechanisms that make them powerful in NLP actually engage on time series data remains unexplored. The persistent competitiveness of simple linear models such as DLinear has fueled ongoing debate, but no mechanistic explanation for this phenomenon has been offered. We address this gap by
A Closed-Form Dual-Barrier CBF Safety Filter for Holonomic Robots on Incrementally Built Occupancy Grid Maps Himanshu Paudel, Basanta Joshi, Dhirendra Raj Madai · 2026-05-06
We present a dual-barrier control barrier function (CBF) safety filter for real-time, safety-critical velocity control of holonomic robots operating in incrementally built occupancy grid maps. As a robot explores an unknown environment, unmapped regions introduce irreducible uncertainty, since obstacle geometry beyond the explored frontier is unknown, making entry into such regions a source of col
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning Lakshita Dodeja, Ondrej Biza, Shivam Vats · 2026-05-06
Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online improvement after demonstrations have been collected. Existing offline-to-online learning methods often cause policies to replace previously learned good actions due to a distribution mismatch between offline data and online learning. In this work, we propose Q2
S-LCG: Structured Linear Congruential Generator-Based Deterministic Algorithm for Search and Optimization Ahmed Qasim Mohammed, Haider Banka, Anamika Singh · 2026-05-06
This study presents a novel deterministic optimization algorithm based on a special variant of the Linear Congruential Generator (LCG). While conventional algorithms generally operate within the search space, the introduced technique follows a two-level architecture. In particular, an external loop that adaptively balances between exploration and exploitation, while the internal loop evaluates sol
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval Nicholas Barnfield, Juno Kim, Eshaan Nichani · 2026-05-06
How many key-value associations can a $d\times d$ linear memory store? We show that the answer depends not only on the $d^2$ degrees of freedom in the memory matrix, but also on the retrieval criterion. In an isotropic Gaussian model for the stored pairs, we show that top-1 retrieval, where every signal must beat its largest distractor, requires the logarithmic model-size scale $d^2\asymp n\log n$
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.