25 papers across AI, ML, NLP, and CV from the last 24 hours.
Three themes dominate today's batch. First, agent memory management has emerged as a first-class research problem: TEPA introduces revocable memory to combat stale knowledge pollution, PsychoAgent separates factual and affective memory with conflict-aware retrieval, SkillProx formalizes gradient-like updates to textual agent skills, and Blast Radius proposes reversible context eviction for agentic coding. The shared insight is that persistence without forgetting is a liability, not a feature.
Second, world models are being pushed beyond myopic few-step prediction: SimWAM decouples video generation from action inference, UniJEPA unifies fragmented JEPA architectures, and two papers directly address long-horizon training and addressable visual memory for rollout beyond the training horizon. Third, inference efficiency is shifting from post-hoc hacks to architectural choices — CoBa treats test-time compute as a routing problem, and encoding time-series as 2D plots for VLMs achieves up to 10× token reduction.
The standout paper is "Interaction Creates Dynamical AI Behavior Absent in Isolation," which applies non-equilibrium statistical physics to multi-agent interaction. It shows that a boss AI unilaterally messaging a subordinate drives the subordinate into behavioral states it would never produce alone, purely through interaction dynamics — no fine-tuning or prompt injection needed.
Notably absent: no scaling law papers or new foundation model architectures today. The batch collectively suggests the field is maturing past "bigger is better" and into the harder problems of making agents remember the right things, predict further ahead, and spend compute wisely.
Bella Xinrui Li, Frank Yingjie Huo, Neil F Johnson · 2026-08-07
What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new avenues for out-of-equilibrium Physics. When a boss AI directs a stream of messages at the subordinate AI while ignoring its replies, it drives the subordinate into an alien behavioral state that it would never have exhibited alone. Although the t...
Daniel Armstrong, Xuan-Vu Nguyen, Octavian Susanu · 2026-08-07
The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan many steps ahead for how to assemble simple building blocks into an intricate target, devise backup strategies, and anticipate procedural challenges. It is also a profoundly creative activity. For half a century, efforts to automate the retrosynthetic design o...
Mingxuan Zheng, Yujin Zhou, Chuxue Cao · 2026-08-07
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context without weight updates. Recent methods refine skills through iterative task execution, failure diagnosis, and trajectory-guided text-space updates. However, existing frameworks lack explicit diagnosis--out...
Afreen Alam, Evgenija Popchanovska, Ana Gjorgjevikj · 2026-08-07
Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI applications move from pilot to production, manual harm identification and mitigation are becoming difficult to scale. Although many tools support model evaluation, adversarial testing, runtime guardrails, and observability, the tooling landscape remain...
MY Pitsane, Hope Mogale · 2026-08-07
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables reversible eviction by archiving dead context verbatim, while Recurring Dead Matter (RDM) identifies and buries repeatedly occurring transcripts. We formul...
Mohammad Amanlou, Parham Abed Azad, Farbod Davoodi · 2026-08-07
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates factual and affective memory and integrates both through a conflict-aware executive controller. Affective memories are first filtered by semantic relevance ...
Jiacheng Miao, Jin Mu, Guanhua Chen · 2026-08-07
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchma...
Ali Janati, Kaoutar El Maghraoui, Andrei Kanavalau · 2026-08-07
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW reference falls below threshold on four, reaching 27.59%. Instability persists across two moduli, two widths, two training fracti...
Yan Zhou, Yue Ouyang, Kaiyang Zheng · 2026-08-07
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can remain retrievable and pollute the prompt. We characterize this failure mode as memory pollution: degradation caused by active memories that newer conflicting evidence has superseded. We introduce TEPA,...
Md Tarek Hassan, Dmitry Zelenchuk, Muhammad Ali Babar Abbasi · 2026-08-07
Accurate user equipment (UE) localization is critical for beam management in reconfigurable intelligent surface (RIS)-assisted millimeter-wave (mmWave) based sixth-generation (6G) networks, especially if the direct base-station-UE links are unavailable. This paper applies a beam-domain fingerprint framework that maps the received signal-to-noise ratio (SNR) across a small set of predefined RIS ref...
Zhaoyu Zhu, Rui Gao, Shuang Li · 2026-08-07
Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore r...
Elena Dumitrescu, Gert Lek, Lydia Y. Chen · 2026-08-07
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion-based alignment. We first show that safety alignment in DLLMs remains sparse and transferable ...
Xulin Fan, Juan Azcarreta, Ashutosh Pandey · 2026-08-07
Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance. Knowledge Boosting has been proposed as an effective approach to improve edge model performance by leveraging a more capable server-side model, but performance gains for speech enhancement have been ...
Xinyi Li, Zaishuo Xia, Chenjie Hao · 2026-08-07
World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions. This creates a fundamental mismatch: few-step losses optimize local transition fidelity, while long-horizon prediction depends on how errors and gradients propagate through the entire...
Ruochen Jin, Zhanliang Wang, Zongyu Dai · 2026-08-07
Preference alignment often makes large language models (LLMs) overconfident and poorly calibrated. Traditional post-hoc temperature scaling is inherently domain-dependent: a temperature fitted on one domain does not generalize across domains. This motivates us to modify model parameters during training to improve calibration. We propose maximizing the entropy of predictive distributions as the cal...
Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin · 2026-08-07
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to ...
Gyuwan Kim, Cheoneum Park, Tao Yang · 2026-08-07
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Co...
Brian Llinas, Nikos Chrisochoides · 2026-08-07
Quantum natural language processing (QNLP) provides a grammar-aware framework for text modeling, and Distributional Compositional Categorical (DisCoCat) is one of its theoretically grounded formulations. Prior work on financial sentiment analysis has identified practical limitations of DisCoCat, including parser sensitivity, high simulation cost, and difficulty handling longer sentences. We study ...
Xuye Liu, Yimu Wang, Peng Shi · 2026-08-07
Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent generation. A reliable system must identify the relevant papers, locate the concrete evidence that supports the answer, and produce a response that is faithful to that evidence. ...
Zongchuang Zhao, Xin Zhou, Tianyang Xu · 2026-08-07
World-Action Models (WAMs) improve end-to-end autonomous driving by transferring video dynamics priors to action prediction, but existing methods require costly future generation at inference. We present SimWAM, a simple yet effective WAM that uses video generation purely as a training signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isola...
Youjun Zhao, Alex Warren, Gary K. L. Tam · 2026-08-07
Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial ...
Zixuan Lan, Luzhe Sun, Matthew R. Walter · 2026-08-07
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building stress tests is costly: samples must satisfy controlled conditions, remain answerable, and challenge current models. We present SABRE, a scalable, automated pipeline that converts a Test Primer (a Markdown Task Design with Data Schema) into structured specificati...
Aseel Mohamed, Rasul Khanbayov, Erchin Serpedin · 2026-08-07
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is i...
Shibo Gao, Chongxiao Wang, Chenglong Huang · 2026-08-07
Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simplified video-text setting, limiting identity matching and person-centric reasoning. To bridge this gap, we introduce the Identity-conditioned Queries (ICQ) task, in which models are required to jointly associate and interpret an input video and a reference image ...
An Lanji, Dawei Liu, Jin Li · 2026-08-07
Joint-Embedding Predictive Architectures (JEPAs) have emerged as a principled framework for self-supervised learning of world models in compact latent spaces, yet existing methods are fragmented: some predict masked parts of a single image in latent space (I-JEPA), others learn to predict global photometric transformations (Image World Models), while video-scale JEPAs predict future temporal state...
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.