arXiv AI Submissions
topic · knowledge/arxiv-ai
§01
about
Newest cs.AI and cs.CL submissions from the arXiv API.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 893 events in this window (995 total on topic). adjust the range or clear it with ALL.
range
01Agentic Permissions Policy Algebra for Taint Confinement in LLM AgentsAutonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon readi{"url":"https://arxiv.org/abs/2607.24625v1","title":"Agentic Permissions Policy …
EVENT. cms4bnpdID. cms4bnpdxa32bkh0c7ojmw70tSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24625v1",
"title": "Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24625",
"version": "1",
"abstract": "Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon reading unvetted data, severely restricting downstream utility. We present APPA (Agentic Permissions Policy Algebra), an IFC framework that resolves this usability bottleneck through engine-managed context branching and prospective acquisition enforcement. Before data acquisition occurs, APPA prospectively evaluates label descents and missing prerequisites, generating actionable remedy plans (Authorize, Accept). To inspect unvetted data without polluting the primary context, a label-seeded child trajectory is spawned, absorbing label descent locally and allowing a trusted sanitizer to return a bounded derivative to the unchanged parent. Governed by a two-monoid model over security labels and shared event logs, we formally prove parent label preservation and merge confinement. Finally, we evaluate APPA on a multi-turn tool-chaining benchmark across four models: it suppresses exfiltration (31%-",
"arxiv_id": "2607.24625",
"categories": [
"cs.CR",
"cs.AI"
],
"published_at": "2026-07-27T16:19:45.000Z"
}02Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature EffectsThe wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended directi{"url":"https://arxiv.org/abs/2607.24645v1","title":"Sparse Autoencoders Encode …
EVENT. cms4bnotID. cms4bnotha329kh0c9hzwcnw7SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24645v1",
"title": "Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24645",
"version": "1",
"abstract": "The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change. Prior work has studied feature geometry inside the model, where features are computed. We instead study the geometry of changes in model logits caused by feature interventions. We introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes. Across SAE variants, consistent one-dimensional effects are rare: few features behave like reusable directions. To interpret this variation, we distinguish value-like features, tied to static information such as factual attributes, from pointer-like features, associated with context-dependent operations. Value-like features more often exhibit structured, low-dimensional effects, although these effects typically span several direct",
"arxiv_id": "2607.24645",
"categories": [
"cs.LG",
"cs.AI",
"cs.CL"
],
"published_at": "2026-07-27T16:45:08.000Z"
}03Efficiency Matters in Autonomous ResearchAI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally importa{"url":"https://arxiv.org/abs/2607.24647v1","title":"Efficiency Matters in Auton…
EVENT. cms4bno8ID. cms4bno86a327kh0czuscvo7cSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24647v1",
"title": "Efficiency Matters in Autonomous Research",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24647",
"version": "1",
"abstract": "AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, is still evaluated primarily by the quality of the final outcome. In this paper, we argue that the efficiency of the solution-search process is an equally important but often overlooked dimension of performance. A strong AR system should not only produce high-quality results, but also reach them with as small a budget as possible. Search efficiency will become increasingly important as AR expands from domains with inexpensive verification, such as mathematics and coding, to real-world scientific settings in which solution evaluation may require costly physical experiments. To capture this dimension, we propose evaluating AR systems using the area under the curve (AUC) of the Pareto frontier, alongside final outcome quality. We compare several families of search algorithms, including hill climbing, beam search, tree search, and evolutionary search, across twelve systems-optimization tasks. We find that no single search structure is consistently the most efficient. We also show that search efficiency and final outcome quality are distinct performan",
"arxiv_id": "2607.24647",
"categories": [
"cs.AI",
"cs.LG"
],
"published_at": "2026-07-27T16:46:33.000Z"
}04Reason-Mediated Behavioral Models for Auditing LLM Social SimulatorsLarge language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-d{"url":"https://arxiv.org/abs/2607.24649v1","title":"Reason-Mediated Behavioral …
EVENT. cms4bnnnID. cms4bnnnua325kh0cgivx2m3xSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24649v1",
"title": "Reason-Mediated Behavioral Models for Auditing LLM Social Simulators",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24649",
"version": "1",
"abstract": "Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-derived reason pattern. We study this problem through a 94-person sunscreen concept test in which each respondent evaluated three product concepts and wrote open-ended rationales. We map those rationales into signed reason states $Z$, where positive signs support adoption and negative signs block it. This gives a practical audit: holding respondent descriptors $D$, category context $K$, and concept treatment $X$ fixed, do human rationale-derived reasons help predict behavior $Y$, and can an LLM simulate the same reason state without seeing the human rationale or outcome? Human rationale-derived reasons substantially improve held-out prediction of purchase intent. LLM-simulated reasons are more brittle: they often sound plausible, but frequently echo the concept board rather than recover the respondent's acceptance or rejection path. The paper contributes an evaluation framework for social",
"arxiv_id": "2607.24649",
"categories": [
"cs.AI"
],
"published_at": "2026-07-27T16:47:27.000Z"
}05Evidence Attribution in Visual Document Understanding without Coordinates or Region LabelsReliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document{"url":"https://arxiv.org/abs/2607.24651v1","title":"Evidence Attribution in Vis…
EVENT. cms4bnmxID. cms4bnmxwa321kh0c03fvevs7SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24651v1",
"title": "Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24651",
"version": "1",
"abstract": "Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investigates whether this failure is partially limited by what the model can express through coordinates. On a verified bilingual CiteVQA subset, we compare the coordinate interface with a language interface in which the model outputs only text, quoting its evidence verbatim, and a multimodal retriever returns the location of each quote as a page region proposed by a layout parser (tables and figures are quoted through their captions or notes); the comparison is repeated over six open vision-language models. Compared with the coordinate interface, evidence recall rises from at most 8 points to between 26 and 47 and the hallucination rate roughly halves, with little change in answer quality. Building ",
"arxiv_id": "2607.24651",
"categories": [
"cs.CV",
"cs.CL",
"cs.IR"
],
"published_at": "2026-07-27T16:49:36.000Z"
}06Kimi K3: Open Frontier IntelligenceWe introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model d{"url":"https://arxiv.org/abs/2607.24653v1","title":"Kimi K3: Open Frontier Inte…
EVENT. cms4bnmdID. cms4bnmdra31zkh0c1jiij8xaSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24653v1",
"title": "Kimi K3: Open Frontier Intelligence",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24653",
"version": "1",
"abstract": "We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model depth. Together with Stable LatentMoE, which effectively activates 16 of 896 routed experts per token, and refined training and data recipes, these advances yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training highlights reinforcement learning across general, agentic, and coding domains and multiple reasoning-effort levels, enabling compositional generalization and robust long-horizon execution. At 2.8T scale, Kimi K3 is supported by infrastructure advances in multiple areas: algorithm-system co-design for KDA, perfectly balanced expert-parallel training with efficient memory management, million-token agentic RL with persistent rollout and sandbox states, and deployment innovations. Extensive evaluations show that Kimi K3 achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and vision tasks. While its o",
"arxiv_id": "2607.24653",
"categories": [
"cs.CL",
"cs.LG"
],
"published_at": "2026-07-27T16:49:54.000Z"
}07A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facilityScientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augme{"url":"https://arxiv.org/abs/2607.24663v1","title":"A corrective agentic hybrid…
EVENT. cms4bnluID. cms4bnluda31xkh0cjj33h3d6SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24663v1",
"title": "A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24663",
"version": "1",
"abstract": "Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augmented Generation, a deployed platform that makes the institutional knowledge at the Advanced Photon Source (APS) accessible to staff through natural-language queries, along with an operations-grounded evaluation. The retrieval engine fuses dense, sparse, and knowledge-graph (KG) channels with query-type-adaptive reciprocal-rank fusion, adds a corrective agentic loop, and runs a native-tool ReAct executor over a Model Context Protocol (MCP) tooling layer. We construct APS-Bench, a 50-question, question-answering (QA) dataset with auditable gold answers. Every retrieval-augmented variant numerically improves strict vital-nugget recall over a naive BM25 baseline (63.8%), with the full corrective Agentic GraphRAG scoring (70.3%). The cross-encoder reranker contributes significantly to answer quality: removing it and allowing the LLM to score relevance drastically reduces strict vital recall b",
"arxiv_id": "2607.24663",
"categories": [
"physics.acc-ph",
"cs.AI",
"cs.IR"
],
"published_at": "2026-07-27T17:01:30.000Z"
}08Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats AccumulatingA language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, w{"url":"https://arxiv.org/abs/2607.24667v1","title":"Eviction as Estimation: A F…
EVENT. cms4bnlaID. cms4bnlaea31vkh0cap5wmt5qSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24667v1",
"title": "Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24667",
"version": "1",
"abstract": "A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, whether an item will be reused, placing existing methods on one axis, the commit lag $H$: online filters and learned predictors commit at $H=0$, while Belady's offline optimum sits where the whole future is known. The missing regime in between, fixed-lag smoothing, waits a bounded number of steps, observes which items a correct near-future prediction attended to, and only then commits. This measurement, demonstrated utility, turns Belady's unobservable future request into something we read off the model itself. We instantiate it as a training-free policy, RMM, a strict generalization of H2O that reduces to it exactly when the measurement is uniform. In controlled settings where reuse is endogenous and separated in time, demonstrated utility identifies used memory far better than accumulated attention, and a small bounded memory behaves like a much larger one. But on independent third-part",
"arxiv_id": "2607.24667",
"categories": [
"cs.AI"
],
"published_at": "2026-07-27T17:08:27.000Z"
}09Co-Learning for Missing Arbitrary Modalities in Multi-modal ClassificationMulti-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training{"url":"https://arxiv.org/abs/2607.24683v1","title":"Co-Learning for Missing Arb…
EVENT. cms4bnkrID. cms4bnkr0a31tkh0cmgv2sboaSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24683v1",
"title": "Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24683",
"version": "1",
"abstract": "Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scenarios subject to operational constraints, such as sensor failures or privacy restrictions, lead to inconsistent modality availability between training and inference times. To handle missing modalities, prior studies have mainly covered bimodal data setups and focused on designing robust fusion processes. Instead, we adopt a multi-modal co-learning framework that prioritizes inter-modal collaboration rather than multi-modal fusion. Specifically, we consider that any subset of modalities may be absent, without assuming predefined missing-modality patterns, an inference scenario we refer to as missing arbitrary modalities. To address this challenge, we introduce two alternative approaches that leverage information at both feature- and decision-level. Experiments on two multi-modal classification benchmarks demonstrate significant robustness gains in various missing modality conditions. The first method shows more robust behavior under minimal missing conditions, where a single modality is absent, whereas the second performs better under ",
"arxiv_id": "2607.24683",
"categories": [
"cs.CV",
"cs.AI",
"cs.LG"
],
"published_at": "2026-07-27T17:23:30.000Z"
}10Beyond Scale and Generation: Understanding Language Model-based Entity MatchingEntity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model varia{"url":"https://arxiv.org/abs/2607.24688v1","title":"Beyond Scale and Generation…
EVENT. cms4bnk7ID. cms4bnk7ra31rkh0cudswdullSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24688v1",
"title": "Beyond Scale and Generation: Understanding Language Model-based Entity Matching",
"source": "arxiv",
"pdf_url": "https://arxiv.org/pdf/2607.24688",
"version": "1",
"abstract": "Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher architectures. However, prior studies often conflate matcher architecture with differences in model backbone, model variant(reflecting different pretraining objectives), and model size, making it difficult to isolate the sources of performance gains. We address this issue through a controlled factorial study spanning three matcher architectures, three model variants and three model sizes from the Qwen3 family, and nine datasets, totaling 1,215 fine-tuning runs. We also evaluate cross-dataset transferability and computational cost. Our results show that model variant is critical for bi-encoders: embedding-oriented variants provide stronger initialization and more favorable representation geometry predictive of downstream matching performance. Cross-encoders retain a consistent advantage over bi-encoders because they jointly encode record pairs rather than representing each record independently, although larger models partially narrow this gap. Generative matchers do not universally outperform cross-encoders",
"arxiv_id": "2607.24688",
"categories": [
"cs.DB",
"cs.CL",
"cs.LG"
],
"published_at": "2026-07-27T17:29:18.000Z"
}showing 1–10 of 893older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above