arXiv Digest — Tuesday, July 14, 2026
25 papers across AI, ML, NLP, and CV from the last 24 hours.
Today's Synthesis
Three clear patterns emerge from today's batch. First, the field is pushing hard on verifiable multimodal reasoning — not just "what is in this video" but "show me the evidence, and preserve it across frames." Papers on evidence-backed video QA, multi-view sports reasoning, and narrative-grounded audio description share the same concern: models need to track entities, relations, and justifications over long horizons, not produce plausible one-off answers. Second, mechanism-level interpretability is maturing past attribution heatmaps. Work on LLM-as-Judge bias at the representation geometry level and an exact instrument for measuring mode usage in state-space models both treat hidden states as measurable objects rather than black boxes. Third, multi-agent systems are surfacing failure modes that don't exist in single-model settings — distributed backdoors where every local monitor passes but the assembled output is harmful.
SpectraReward stands out for its simplicity. Rather than training a reward model for text-to-image generation, it uses a pretrained MLLM's own prompt-recovery log-likelihood as an off-the-shelf reward signal. The elegance is that it requires zero additional training, no human preference data, and no decomposed verification questions — just one teacher-forced forward pass.
Notably absent are papers on pure LLM scaling or architecture novelties. The batch leans toward understanding, grounding, and compositional safety rather than raw capability gains. That shift suggests the field is entering a phase where how models reason, where they fail in composition, and whether their outputs can be independently verified matters more than the next benchmark point.
Papers by Category
Computer Vision
Runhui Huang, Qihui Zhang, Zhe Liu · 2026-07-13
In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-condi...
Daniel Garibi, Ronen Kamenetsky, Hadar Averbuch-Elor · 2026-07-13
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits. We present a method for fine-grained identity tuning in text-to-image personalization mo...
Shijie Wang, Honglu Zhou, Ziyang Wang · 2026-07-13
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual answers without verifiable visual grounding. Existing explainability efforts rely on textual rationales or sparse bounding boxes, which struggle to capture complex video dynamics such as occlusions and non-rigid deformations. We propose Evidence-Backed Video Quest...
Kerui Chen, Jinglu Wang, Xiaoyi Zhang · 2026-07-13
Recent Multimodal Large Language Models (MLLMs) achieve strong performance on single-view video understanding benchmarks. However, sports videos involve dense occlusion, rapid motion, and complex interactions that are difficult to resolve from a single viewpoint. In practice, sports events are recorded from multiple camera angles, providing complementary evidence used by referees. Yet, no existing...
Caleb Robinson, Anthony Ortiz, Simone Fobi Nsutezo · 2026-07-13
When a large disaster strikes, responders need a map of which buildings are damaged within hours. The models that do well on public benchmarks assume matched before-and-after imagery and a training set drawn from similar past events, and neither is usually available for a new disaster in its first day. We present HASTE (High-speed Assessment and Satellite Tracking for Emergencies), a no-code web p...
Zihan Su, Teng Hu, Jiangning Zhang · 2026-07-13
Autoregressive diffusion models have enabled high-quality video generation, yet their sequential nature inherently suffers from error accumulation. In long-horizon video synthesis, minor prediction deviations compound over time, inevitably leading to unconstrained generative drift, structural collapse, and severe visual degradation. To address this, we propose Cycle-World, a novel framework design...
Kaixin Ma, Di Feng, Alexander Metz · 2026-07-13
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goa...
Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin · 2026-07-13
Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and story context across scenes so that blind and low-vision (BLV) audiences can follow a film. Modern video-language models (VLMs) are effective on short clips, but they often treat each moment independently, producing descriptions that miss who characters are, why e...
Yilong Yang, Jianxin Tian, Shengchuan Zhang · 2026-07-13
Referring Camouflaged Object Detection (Ref-COD) requires segmenting hidden targets guided by reference cues. While supervised methods are annotation-heavy and training-free approaches via sparse point-prompting are sensitive to localization errors, we propose GFR-SAM, a robust three-stage training-free framework. GFR-SAM shifts the paradigm from fragile point-matching to a "Generate-Filter-Refine...
Computation & Language (NLP)
Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu · 2026-07-13
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication, and more. In recent years, it has become increasingly recognized as a cornerstone of capable, transparent AI systems. Yet while LLMs have made significant progress across diverse real-world tasks, it is not yet clear when, how, or to what extent they can exhibi...
Zixiang Xu, Sixian Li, Huaxing Liu · 2026-07-13
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operationally useful in ways it does not afford. We report three findings, across seven ...
Lingkai Kong, Zijian Wu, Yuzhe Gu · 2026-07-13
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of ...
Elmira Salari, Hazem Amamou, Jose Victor de Souza · 2026-07-13
Retrieval-Augmented Generation (RAG) has been increasingly adopted to reduce hallucinations and strengthen the factual grounding of large language models (LLMs). While robustness to errors in the retrieval process has been explored, the impact of ideological bias on LLM outputs has been overlooked. For instance, if the retrieved material contains ideological positions, the RAG may transmit, amplif...
Ayoung Lee, Ryan Kwon, Yunxiang Zhang · 2026-07-13
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory...
Iman Johary, Guillaume Bied, Alexandru C. Mara · 2026-07-13
Career paths encode decades of skill acquisition, role transitions, and educational investment, and understanding them at scale underpins workforce planning, labor market policy, and job recommendation. Resumes are a rich source of information about career paths: they contain detailed descriptions of work experience, education, and skills. Yet their unstructured, heterogeneous, and multilingual na...
Machine Learning
Tiberiu Musat, Tiago Pimentel, Nicholas Zucchet · 2026-07-13
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning. In this c...
Bijan Mazaheri, Jiaqi Zhang, Caroline Uhler · 2026-07-13
Causal discovery algorithms learn a network that describes the causal dependencies among random variables. A common workflow involves first utilizing conditional independence properties on observational data to determine partially directed causal relationships, then applying interventions to orient the unknown causal directions. A critical assumption for the first step is faithfulness: a requireme...
Raktim Bhattacharya · 2026-07-13
Selective state-space models such as Mamba route information through a bank of first-order modes whose input coupling is set by a learned selection mechanism. We give an exact instrument for measuring how a trained model uses these modes. Because the state matrix is diagonal, each channel's output decomposes exactly into per-mode contributions, and a per-(layer, channel, window) Gram tensor yields...
Haozhe Huang, Yudong Xu, Abhijoy Mandal · 2026-07-13
Discrete diffusion models offer a powerful framework for solving complex reasoning tasks, particularly through compositional generation, which combines multiple pre-trained experts to generalize beyond their individual training data. Recent theoretical corrections introduce time-dependent mixing weights to better align composed diffusion dynamics with the intended target. However, these methods ar...
Shambhavi Balamuthu Sampath, Behzad Shomali, Nael Fasfous · 2026-07-13
With deep neural networks (DNNs) increasingly deployed on edge devices, hardware (HW)-aware optimization techniques--such as HW-aware compression and HW-aware neural architecture search (HW-NAS)--have become essential. These methods rely on real feedback from the target hardware to tailor DNN architectures for efficient deployment. While the search can be parallelized, latency measurements via har...
Alper Kamil Bozkurt, Shangtong Zhang, Yuichi Motai · 2026-07-13
Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subsequently improved through limited online interaction. This offline-to-online RL (O2O-RL) paradigm is particularly promising in nonstationary domains where interaction is costly or potentially hazardous. Standard O2O-RL pipelines train multiple candidate policie...
Robotics
Yunhai Feng, Natalie Leung, Jiaxuan Wang · 2026-07-13
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and...
Zhiyang Dou, John U. Onyemelukwe, Hangxing Zhang · 2026-07-13
Differentiable simulators have advanced policy learning and model-based control, yet actuator dynamics remain an important source of sim-to-real error. This is particularly acute on low-cost platforms, where the linear current-to-torque relation tau= K_tI becomes unreliable during commanded-target tracking because of friction, hysteresis, backlash, and thermal effects. We present NeuralActuator, a...
Artificial Intelligence
Ziv Ben-Zion, Teddy Lazebnik · 2026-07-13
Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes how recipients actually behave, and through what channel, remains unknown. Here, in a randomized crossover field experiment, 121 employees across six companies sent work emails under three conditions over three weeks: unaided writing, GPT-5 rewriting in a playful tone, and GPT-5 rewr...
Cryptography & Security
Yibo Hu, Ren Wang · 2026-07-13
As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The probl...
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.