arXiv AI Submissions

topic · knowledge/arxiv-ai
DOC.
knowledge/arxiv-ai
REV.
995 evt
DATE.
05-JUN-2026
SCOPE.
custom
§01

about

Newest cs.AI and cs.CL submissions from the arXiv API.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 902 events in this window (995 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in SportsLarge language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesiz{"url":"https://arxiv.org/abs/2607.24573v1","title":"LLM-SoccerArena: Benchmarki…
EVENT. cms4bnuaID. cms4bnuaja32tkh0colgoyw3hSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24573v1",
  "title": "LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24573",
  "version": "1",
  "abstract": "Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLM",
  "arxiv_id": "2607.24573",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:38:29.000Z"
}
02CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video UnderstandingLong-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides{"url":"https://arxiv.org/abs/2607.24582v1","title":"CADER: Confidence-Aware Dyn…
EVENT. cms4bntrID. cms4bntr2a32rkh0chn5qyxdgSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24582v1",
  "title": "CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24582",
  "version": "1",
  "abstract": "Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example regardless of difficulty. This uniform strategy invokes unnecessary tool-assisted processing for easy questions and provides limited control when difficult questions require fine-grained temporal evidence. We propose CADER (Confidence-Aware Dynamic Evidence Reasoning), a training-free framework for adaptive and reliable long-video reasoning. CADER first performs global reasoning over uniformly sampled frames and estimates answer confidence with a logit-margin signal, allowing high-confidence examples to exit early. For uncertain examples, CADER activates a second-stage tool-augmented loop that combines temporal cropping, lightweight semantic verification, and Relevance-Guided Resampling to progressively localize question-relevant evidence. This design treats tool use as a sample-level decision: a single global pass handles easy cases, while additional reasoning is reserved for examples where uncertainty suggests that more evidence is needed. Experiments on multiple VideoQA benchmarks show that CADER improves ",
  "arxiv_id": "2607.24582",
  "categories": [
    "cs.CV",
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:49:42.000Z"
}
03From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile InferenceWe present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We deve{"url":"https://arxiv.org/abs/2607.24585v1","title":"From Data to Device: ELMOD …
EVENT. cms4bnt7ID. cms4bnt7ka32pkh0c95e16hsySRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24585v1",
  "title": "From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24585",
  "version": "1",
  "abstract": "We present ELMOD - Efficient Language Model for On-Device Deployment - a compact (2.7B) German language model designed for efficient inference on resource-constrained hardware. ELMOD was trained on a limited computational budget (55k H100 GPU hours) using exclusively publicly available data. We developed a suite of German-specific data pre-processing, which differ from English-oriented counterparts in their handling of morphological variation, compounding, and orthographic conventions. Furthermore, we introduced a quality filtering and rephrasing step, which increased the instructional quality of the data, improved performance during the annealing phase, and reduced overall compute requirements. Thanks to our architectural model and data choices, including prefiltering, our educational-quality filtering and rephrasal to raise the educational-quality, ELMOD is the strongest performer in its size class (<3B), matching the performance of 7B-parameter models in German.",
  "arxiv_id": "2607.24585",
  "categories": [
    "cs.CL"
  ],
  "published_at": "2026-07-27T15:51:41.000Z"
}
04D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language ModelsLarge Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple sp{"url":"https://arxiv.org/abs/2607.24586v1","title":"D-Score: A Spectral Hidden-…
EVENT. cms4bnsnID. cms4bnsnta32nkh0ce7v0zxi2SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24586v1",
  "title": "D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24586",
  "version": "1",
  "abstract": "Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity. The motivation is that, when a model processes a text that conflicts with information available in its own internal state, the hidden representation may encode both the asserted content and some form of counter-evidence, uncertainty, correction, or lack of support; this can make the hidden trajectory spread across additional singular directions. We formalize this intuition through a lightweight spectral argument and evaluate the resulting detector on FAVA-Annotation and RAGTruth. The experiments indicate t",
  "arxiv_id": "2607.24586",
  "categories": [
    "cs.CL",
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:52:33.000Z"
}
05SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM AgentsEarly warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent adv{"url":"https://arxiv.org/abs/2607.24588v1","title":"SIREN: Towards End-to-End E…
EVENT. cms4bns4ID. cms4bns4ka32lkh0carh9e0bkSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24588v1",
  "title": "SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24588",
  "version": "1",
  "abstract": "Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events. However, expert-centered warning workflows are costly, labor-intensive, and difficult to scale throughout the warning-to-action process. Although recent advances in Large Language Model (LLM) agents have enabled the automation of weather-related tasks, existing studies remain centered on isolated scientific tasks and overlook the chain of interdependent processes required for operational extreme-weather early warning. To bridge this gap, this study investigates automated end-to-end extreme-weather early warning through LLM agents. We first develop SIREN-Bench, a comprehensive benchmark comprising 600 question-answer instances across 19 tasks, and covering four individual warning procedures and an end-to-end warning chain. Evaluation on SIREN-Bench reveals substantial capability gaps in existing weather agent frameworks. This motivates us to develop SIREN, an experience-grounded agent framework inspired by experts' use of historical cases, which combines an agentic execution environment integrating heterogeneous weather evidence and tools wi",
  "arxiv_id": "2607.24588",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:53:19.000Z"
}
06Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future DirectionsThe development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Explo{"url":"https://arxiv.org/abs/2607.24589v1","title":"Artificial Intelligence and…
EVENT. cms4bnrlID. cms4bnrl1a32jkh0c7jo0r9x2SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24589v1",
  "title": "Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24589",
  "version": "1",
  "abstract": "The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared achievements. The rise of Artificial Intelligence (AI) has significantly accelerated the global processes of digitization, informatization, and intelligence. Exploring how AI can leverage inherent characteristics to influence the development trajectory of IE is a topic that warrants further investigation. Given AI's increasing prominence and role within IE, the paper analyzes this new form, examining both AI's unique contributions to IE and its potential challenges. Firstly, the paper synthesizes the conceptual frameworks surrounding IE, decomposing them into manifestations in physical, social, and thinking spaces. Furthermore, the concept of Artificial Intelligence IE (AIIE) is introduced from a spatial perspective, with an exploration of the characteristics AI contributes to IE. Subsequently, the paper employs an evolutionary perspective to analyze the roles provided by AI during different development periods of AIIE. The paper then verifies the feasibility, effectiveness, and rationality of the AIIE's definition and analyzes AIIE development fr",
  "arxiv_id": "2607.24589",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:54:23.000Z"
}
07PIVOT: Efficient Query-Group Indexing for Token-Level Sparse AttentionToken-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incur{"url":"https://arxiv.org/abs/2607.24593v1","title":"PIVOT: Efficient Query-Grou…
EVENT. cms4bnr1ID. cms4bnr19a32hkh0c4fx5y0hzSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24593v1",
  "title": "PIVOT: Efficient Query-Group Indexing for Token-Level Sparse Attention",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24593",
  "version": "1",
  "abstract": "Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incurring a cost of O(L^2) per layer for a sequence of length L. We observe that this per-query scan is largely redundant: nearby queries select highly overlapping top-k tokens, and the indexer scores are long-tailed along the key axis. We exploit these properties in PIVOT, Proxy Indexing Via One full-prefix Traversal, a training-free, drop-in replacement for the DSA indexer that shares one prefix scan across a group of nearby queries. PIVOT aggregates a group into a single proxy query, performs one shared full-prefix scan to obtain a candidate set, and then selects a top-k for each query from that set. Two variants trade speed for fidelity: PIVOT-Reuse shares the proxy top-k across the group for maximum speed, whereas PIVOT-Refine re-scores the candidate set with the indexer of each query and then selects an individual top-k, matching the dense indexer at a small additional cost. A single al",
  "arxiv_id": "2607.24593",
  "categories": [
    "cs.CL"
  ],
  "published_at": "2026-07-27T15:58:07.000Z"
}
08Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code ReviewBackground: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them. The role of {"url":"https://arxiv.org/abs/2607.24601v1","title":"Evaluating the Impact of Ex…
EVENT. cms4bnqhID. cms4bnqh4a32fkh0cp49mddvbSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24601v1",
  "title": "Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24601",
  "version": "1",
  "abstract": "Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains hard to understand. Developers struggle to assess the validity of LLM-generated reviews, making it difficult to gauge how much trust to place in them. The role of Explainable AI (XAI) in code review and its impact on trust remain underexplored. Objective: We study the influence of XAI on developer trust in AI-assisted code reviews. Method: We conducted a within-subjects user study with 34 participants, comparing three LLM-based code review systems with varying levels of XAI support: Condition A (detailed explanation and review feedback), Condition B (review feedback only), and Condition C (no explanations). Participants reviewed real-world code change requests alongside the AI-generated reviews. We measured trust perceptions, agreement with the AI recommendation, the reasoning given for each decision, and the time taken. Results: The level of explanation significantly influences both trust and agreement with AI recommendations, but in different ways. Full explanations (A) yield the highest perceived trust (M = 3.99/5) but not the highest agreement",
  "arxiv_id": "2607.24601",
  "categories": [
    "cs.SE",
    "cs.AI",
    "cs.HC"
  ],
  "published_at": "2026-07-27T16:04:24.000Z"
}
09Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code RepairGenerate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories.{"url":"https://arxiv.org/abs/2607.24604v1","title":"Looping Is Not Reliability:…
EVENT. cms4bnpxID. cms4bnpxja32dkh0cjpui99zjSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24604v1",
  "title": "Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24604",
  "version": "1",
  "abstract": "Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95\\% CI $[8.9,37.0]$, exact Holm $p=0.0337$). A prospective 540-rollout policy eliminates observed correct-start harm but reduces wrong-start repair and fails its joint criterion. Repository experiments over 24 bugs and four coder stacks expose floor effects and component heterogeneity without Holm-significant effects. We therefore separate admission, preservation, grounded certification, competence, and liveness. We derive an evidence-bound typed loop contract and instantiate ",
  "arxiv_id": "2607.24604",
  "categories": [
    "cs.CL",
    "cs.AI"
  ],
  "published_at": "2026-07-27T16:05:23.000Z"
}
10Agentic Permissions Policy Algebra for Taint Confinement in LLM AgentsAutonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon readi{"url":"https://arxiv.org/abs/2607.24625v1","title":"Agentic Permissions Policy …
EVENT. cms4bnpdID. cms4bnpdxa32bkh0c7ojmw70tSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24625v1",
  "title": "Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24625",
  "version": "1",
  "abstract": "Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon reading unvetted data, severely restricting downstream utility. We present APPA (Agentic Permissions Policy Algebra), an IFC framework that resolves this usability bottleneck through engine-managed context branching and prospective acquisition enforcement. Before data acquisition occurs, APPA prospectively evaluates label descents and missing prerequisites, generating actionable remedy plans (Authorize, Accept). To inspect unvetted data without polluting the primary context, a label-seeded child trajectory is spawned, absorbing label descent locally and allowing a trusted sanitizer to return a bounded derivative to the unchanged parent. Governed by a two-monoid model over security labels and shared event logs, we formally prove parent label preservation and merge confinement. Finally, we evaluate APPA on a multi-turn tool-chaining benchmark across four models: it suppresses exfiltration (31%-",
  "arxiv_id": "2607.24625",
  "categories": [
    "cs.CR",
    "cs.AI"
  ],
  "published_at": "2026-07-27T16:19:45.000Z"
}
showing 1–10 of 902older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above