40 papers across AI, ML, NLP, and CV from the last 24 hours.
Three themes cut across today's batch. First, mechanistic understanding is replacing surface-level metrics: papers on SAE feature instability, attention-path fragility as uncertainty, and probabilistic consistency verification all attack the problem of how do we actually know what a model knows — not by probing outputs, but by examining internal structure. Second, vision systems are moving from closed benchmarks to open-world reasoning: camouflaged object detection drops its closed-world assumption, 3D Gaussian splatting takes on hierarchical commonsense reasoning, and GUI agents learn to self-evolve at test time rather than freezing after deployment. Third, real-world deployment constraints are driving methodology: loss-resilient compression for satellite links, corpus-free unlearning for privacy compliance, and agentic coding's "catastrophic remembering" all stem from systems hitting production edges.
The standout is Kushal Chakrabarti's "Why Does CLAUDE.md Keep Growing?" — it names a phenomenon every AI-assisted developer has watched happen in real time, then formalizes it as catastrophic remembering, the inverse of catastrophic forgetting. The argument that deleting an instruction without its rationale costs exponential effort in prompt size is both simple and genuinely alarming for repository hygiene.
Notable by absence: almost nothing on scaling laws or new base model architectures today. The field is looking inward — at interpretability, robustness, and deployment behavior — rather than outward at bigger models. That suggests a maturing phase where the questions are shifting from "how large can we go" to "how well do we understand what we've built."
Mingju Gao, Jingkai Zhou, Kun Gai · 2026-05-29
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-matching losses. However, directly optimizing Fréchet objectives can cause Fréchet hacking. The target metrics keep improving, but visual quality and Fréchet alignment in other feature spaces may stagnate or deteriorate. W...
Wenrui Bao, Tianyun Jiang, Zhiben Chen · 2026-05-29
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, and bimanual coordination. Endoscopic video is comparatively inexpensive and abundant relative to ...
Yizhou Xu, Lars Bretzner, Tiesheng Wang · 2026-05-29
This paper presents a self-supervised representation learning framework for understanding 3D skeleton-based human motion in soccer, using future motion prediction as the learning objective. Since human motion is inherently uncertain, accounting for multiple plausible futures is essential for capturing the underlying motion dynamics and learning effective representations. To this end, we introdu...
Bowei Liu, Zheng Lu, Yuhan Bian · 2026-05-29
Recent advances in video generation models have significantly improved the realism of synthetic videos, blurring the boundary between generated and authentic content and raising concerns about misinformation. Existing MLLM-based detectors mainly rely on supervised fine-tuning or label-level reinforcement learning, where coarse supervision limits generalization to unseen scenarios and emerging v...
Jiayu Ding, Meilu Song, Yun Chen · 2026-05-29
While 3D Gaussian Splatting (3DGS) has advanced open vocabulary scene understanding, existing methods remain confined to explicit queries. They struggle to interpret implicit intents, complex spatial constraints, and commonsense reasoning required for practical embodied interactions. To address this gap, we introduce the task of reasoning 3D Gaussian segmentation and construct two benchmarks, C...
Huafeng Chen, Yueming Lyu, Ziyuan Chen · 2026-05-29
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in storing and recalling rich person-related knowledge, raising increasing concerns about reliable knowledge removal. However, existing machine unlearning approaches for MLLMs typically assume access to original forget and retain corpora, which are often unavailable in realistic deletion scenarios. To address thi...
Moti Rattan Gupta, Anupam Sobti · 2026-05-29
Agricultural monitoring faces unique challenges, arising from the landscape's complex temporal, phenological, and climate dynamics, yet monitoring them is critical for ensuring food security. Synthetic Aperture Radar (SAR) satellites offer all-weather day-night imaging capability supporting key monitoring tasks including crop type mapping, yield prediction and phenological event detection. Exis...
Huafeng Chen, Yueming Lyu, Chenyang Si · 2026-05-29
Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in recent years. However, most existing COD methods are developed under a closed-world assumption, where each input image is assumed to contain a camouflaged object. This assumption ignores realistic scenarios with pure backgrounds or non-camouflage...
Yuhang Wei, Chuqin Zhou, Yibo Shi · 2026-05-29
Learned image compression (LIC) has achieved impressive rate-distortion performance. However, existing methods remain highly vulnerable to packet loss, a common challenge in satellite and emergency communications. This vulnerability stems from non-uniform information distribution at the packetization stage and sequential decoding dependencies at the entropy coding stage. We propose an end-to-en...
Hang Li, Jiahe Li, Meiying Gu · 2026-05-29
Feed-forward Gaussian reconstruction has recently emerged as an efficient approach for driving scene reconstruction. However, prevailing LiDAR-based methods preserve the initial correspondence between observed points and Gaussian primitives, treating the initialized primitive set as the final representation. Unlike optimization-based 3DGS, these methods cannot accumulate gradients during traini...
Biratal Raj Wagle, Bashirul Azam Biswas, Grant Chau · 2026-05-29
Automated lesion segmentation in whole-body PET/CT imaging can assist clinicians with cancer detection, staging, and treatment planning across radiotracers and cancer types. However, training lesion segmentation models that capture variations in lesion size, distribution, and appearance requires large annotated datasets, whose creation is both time- and expertise-intensive. We propose FEEDS, a ...
Hesam Araghi, Jan van Gemert, Nergis Tomen · 2026-05-29
Event cameras capture intensity changes asynchronously with high temporal resolution, requiring novel preprocessing methods for downstream tasks. Unlike static intensity snapshots, event data inherently encode information about scene dynamics and object motion, meaning that features derived from events can exhibit behaviors with no direct analogue in frame-based vision. In this paper, we analyz...
Mouxiao Huang, Qiangyu Yan, Borui Jiang · 2026-05-29
Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (e.g., CIDEr and SPICE) and LLM-as-scorer protocols struggle to verify dense factual claims, while existing QA-based alternatives generally offer lower probe density, narrower domain coverage, or no explicit alignment between individual questions...
Ali Saleh, Abdul Karim Gizzini, Mohamad Ghassany · 2026-05-29
Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and prediction. Consequently, in critical domains such as remote sensing, where high-resolution imagery mus...
Ismail Ismail Tijjani, Sunusi Muhammad Ibrahim, Amina Ibrahim Khaleel · 2026-05-29
The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches rely on controlled datasets that do not adequately represent real-world farming conditions, particularly in underrepresented regions such as Africa. This study presents a comparative evaluation of six object detection models -- ...
Chen Lyu, Xingwei Tan, Simon Cullen · 2026-05-29
Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversat...
Nikolai Bolik, Lennart Stöppler, Artur Andrzejak · 2026-05-29
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level mea...
Orr Paradise, Oliver Richardson, Yoshua Bengio · 2026-05-29
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially caused by an AI action. We construct an interactive PCP as follows. Let a predictive model be spe...
Eric A. F. Reinhardt, Adam J. Hauser · 2026-05-29
The attention mechanism forms the foundation of many modern AI models such as the Transformer. In one subclass of problems where attention is used, inputs and outputs are bound to the probability simplex so that all outputs sum to one. In this setting, softmax attention admits an exact, component-by-component quantum realization. Attention scores are Hadamard-test statistics on block-encoded pr...
Changhao Xiang, Shangyu Xing, Zhen Wu · 2026-05-29
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, lea...
Nguyen Thai Anh, Truong Viet Vu, Tran Thien Thanh · 2026-05-29
The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the m-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data. We propose hierarchical empirica...
Pavel Averin, Theodoros Moysiadis, Ioannis Katakis · 2026-05-29
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI decisions. This survey reviews CI testing with emphasis on assumptions, robustness, and scalability in high-dimensional and mixed-type settings common in biomedic...
Shiqi Huang, Jiani He, Dingyan Shang · 2026-05-29
Detecting or attributing a supply-chain disruption is not the same as selecting the intervention that maximizes recoverable net value. We present CriticalSCM-Bench v1, a controlled synthetic benchmark with causal ground truth, paired factual/counterfactual rollouts, and an explicit net-value objective. Relative to a full-information train-selected static benchmark, LambdaMART improves median no...
Zetao Hong, Song Yuan, Yuanhao Ding · 2026-05-29
Modern reinforcement learning (RL) post-training pipelines for large language models (LLMs) increasingly combine rollout workloads across multiple domains and feedback paradigms. Prefix-aware routing improves inference efficiency through cache reuse and load balancing, but it does not control how heterogeneous rollout sessions compete for KV-cache capacity. When reinforcement learning with veri...
Kaan Buyukkalayci, Kyle Pak, Merve Karakas · 2026-05-29
This paper develops a method for sensor-subset selection for tracking. Prior work showed that low-cost acoustic Received Signal Strength Indicator (RSSI) measurements can be used to recommend subsets of sensor nodes whose expensive sensing modalities, such as cameras, can achieve high tracking accuracy. While efficient, RSSI-based approaches are challenged by acoustic interference. We propose a...
Vladimir Iglovikov · 2026-05-29
Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. Code paths that choose these values separately can silently misalign the data. AlbumentationsX keeps the transform list, probabilities, annotation settings, and random se...
Kiran Madhusudhanan, Christian Klöttergens, Lars Schmidt-Thieme · 2026-05-29
Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existing approaches often face a fundamental trade-off between distributional flexibility and accurate mean prediction. Traditional parametric methods, such as Mean Variance Estimation (MVE), can suffer from degraded point accuracy when trained under joint Negativ...
Kushal Chakrabarti · 2026-05-29
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D| instructions. We name the r...
Songlin Du, Xiaoyong Lu, Zeyu Wu · 2026-05-29
Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain hig...
Robert Bitterling, Christian Nettersheim, Jörn Hees · 2026-05-29
Low Power Wide Area Networks like LoRa are increasingly deployed for smart city applications, requiring accurate path loss prediction for effective network planning. Traditional (empirical) propagation models often exhibit limited accuracy in these scenarios. We investigate machine learning models for LoRa path loss prediction, systematically analyzing how prediction accuracy scales with traini...
Artyom Sabitov, Daniil Volkov, Alexey Zaytsev · 2026-05-29
Large-scale neural recommender systems are typically trained with a softmax cross-entropy objective over the full item vocabulary. For a typical large number of possible items K, the final classification layer dominates memory, requiring O(nK) logits and gradients to materialize for a batch of n examples. Sampled softmax reduces this cost by restricting the objective to only k << K candidate ne...
Sepideh Saran, Mahsa Ghanbari, Uwe Ohler · 2026-05-29
Deep learning models have emerged as the standard computational tool for a wide range of applications in genomics. Yet, uncertainty quantification (UQ) -- and more specifically, the reliability of different uncertainty estimates in this domain -- has received little systematic attention. This work presents an empirical analysis of UQ in deep learning models, focusing on genomics applications. I...
Nikita Sevriukov, Anna Barabanova, Uliana Gagarina · 2026-05-29
Inverse reinforcement learning (IRL) aims to recover a reward function under which the resulting policy reproduces the behavior observed in expert demonstrations. A natural approach is to formulate IRL as a bilevel optimization problem, in which the inner level corresponds to policy optimization under the learned reward and the outer level measures the discrepancy between the induced policy and...
Shiyu Xuan, Zechao Li · 2026-05-29
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon failed exploration. To overcome this, we propose a Test-Time Self-Evolving framework that enables m...
Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle · 2026-05-29
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, clas...
Abigail Oppong, P Sam Sahil, Tadesse Destaw Belay · 2026-05-29
Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that p...
Minsoo Kim, Sungyoung Ji, Kisung Moon · 2026-05-29
We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is fragile under perturbation of its attention pathways. We instantiate this as ASMI (Attention-Subnetwork Mutual Information), a training-free estimator that masks attention heads and measures the BALD mutual information among the result...
Sourabrata Mukherjee, Kalika Bali, Sunayana Sitaram · 2026-05-29
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet those actions are the product: they fix cost and latency, decide how the system fails, and are the only auditable part of its behaviour. We make the action policy the measured object across 8 model...
Ebrahim Khaled Ebrahim · 2026-05-29
Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are sparse -- a common issue with continuous predictors like age or weight, where the asymptotic distributional assumptions are not satisfied. This t...
Emanuele Dolera, Stefano Favaro, Matteo Giordano · 2026-05-29
We study posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. We embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link condition on the Fisher information and suitable local regularity assumptions, w...
This digest is generated automatically from arXiv submissions. Not affiliated with arXiv or Cornell University.