arXiv AI Submissions

topic · knowledge/arxiv-ai
DOC.
knowledge/arxiv-ai
REV.
995 evt
DATE.
05-JUN-2026
SCOPE.
custom
§01

about

Newest cs.AI and cs.CL submissions from the arXiv API.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 929 events in this window (995 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01Penelope: Localized Latent Recurrence for Efficient Structured ReasoningComplex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning {"url":"https://arxiv.org/abs/2607.25915v1","title":"Penelope: Localized Latent …
EVENT. cms5r49aID. cms5r49agagjrkh0c1g0l5fdgSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25915v1",
  "title": "Penelope: Localized Latent Recurrence for Efficient Structured Reasoning",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25915",
  "version": "1",
  "abstract": "Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while redu",
  "arxiv_id": "2607.25915",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:06:46.000Z"
}
02Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QAIn this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annot{"url":"https://arxiv.org/abs/2607.25921v1","title":"Evaluating VLMs for Autonom…
EVENT. cms5r48qID. cms5r48qhagjpkh0cbbfycc07SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25921v1",
  "title": "Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25921",
  "version": "1",
  "abstract": "In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeline focusing on geometry clipping. In this evaluation, a custom exploration agent navigates a game level to collect visual observations, while the automatic annotation pipeline provides frame-level clipping labels. This setup allows us to evaluate recent VLMs on a controlled anomaly detection task without manual annotation. We benchmark six recent VLMs (Gemini, GPT, Qwen, Gemma, Llama, and Ministral) under a zero-shot prompting setting and analyse their sensitivity to four prompt variants. Our results show that while the VLMs can capture visual cues associated with geometry clipping, they all produce substantial false positives on visually ambiguous frames such as near-contact geometry and partial occlusions. Gemini-3.1-Flash achieves the best overall accuracy and is the most robust to prompt variation, while open-source models exhibit large precision--recall swings depending on the prompt design. These findings suggest that current VLMs are best suited as high-recall candidate filters within multi-stage QA pipelines rather than as standalone bug",
  "arxiv_id": "2607.25921",
  "categories": [
    "cs.CV",
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:10:47.000Z"
}
03dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision TreesOver the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a current state-of-the-art tool. However, for systems that are large or have many corner cases, even such representations tend to be too complex and not human-comp{"url":"https://arxiv.org/abs/2607.25925v1","title":"dtControl2+$\\varepsilon$: …
EVENT. cms5r486ID. cms5r486oagjnkh0cgdop91jeSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25925v1",
  "title": "dtControl2+$\\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25925",
  "version": "1",
  "abstract": "Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a current state-of-the-art tool. However, for systems that are large or have many corner cases, even such representations tend to be too complex and not human-comprehensible. Unfortunately, reducing the size of the decision tree is not straightforward, as missing just a single crucial case might result in an incorrect controller. We tackle this issue in the setting of Markov decision processes, extending dtControl2 by \"$\\varepsilon$\" functionality: Given an allowed imprecision $\\varepsilon \\geq 0$, we construct a smaller decision tree, distilling the essence of the controller, while still guaranteeing its $\\varepsilon$-optimality. This enables us to provide tunably simpler explanations, omitting a controllable amount of detail. Our tool constructs decision trees that are orders of magnitude smaller than the state of the art.",
  "arxiv_id": "2607.25925",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:15:11.000Z"
}
04Face De-Identification: A Domain-Centric Survey from Capture to ProcessingFace de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition while preserving utility for downstream tasks. With the rising emphasis on data privacy and responsible AI, face De-ID has emerged as an active researc{"url":"https://arxiv.org/abs/2607.25926v1","title":"Face De-Identification: A D…
EVENT. cms5r47mID. cms5r47m8agjlkh0c5lw8ylkfSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25926v1",
  "title": "Face De-Identification: A Domain-Centric Survey from Capture to Processing",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25926",
  "version": "1",
  "abstract": "Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity recognition while preserving utility for downstream tasks. With the rising emphasis on data privacy and responsible AI, face De-ID has emerged as an active research area spanning computer vision and privacy-preserving communities. Early approaches, and many contemporary ones, operate in the digital domain by modifying pixel-level or appearance-level features through post-capture processing. Recent advances extend face De-ID beyond post-processing by integrating privacy mechanisms directly into sensors during image acquisition, bridging sensing systems and downstream vision algorithms. In parallel, physical-domain methods explore wearable accessories and materials that conceal identity information in real-world environments prior to capture. In this survey, we present the first unified overview that spans the full data acquisition pipeline, encompassing the physical, sensor, and digital domains. Through this domain-centric lens, we systematically analyze current methodologies, technical progress, and the distinct challenges inherent to each stage. ",
  "arxiv_id": "2607.25926",
  "categories": [
    "cs.CV",
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:15:12.000Z"
}
05Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical CasesClinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reason{"url":"https://arxiv.org/abs/2607.25933v1","title":"Evaluating Multi-Turn Multi…
EVENT. cms5r472ID. cms5r472fagjjkh0chgqf03tzSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25933v1",
  "title": "Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25933",
  "version": "1",
  "abstract": "Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning. However, existing evaluations of multimodal large language models (MLLMs) typically rely on single-turn or isolated tasks, making it difficult to fully capture the complexity of real-world clinical diagnosis. To bridge this gap, we developed ClinMM-Bench, the largest multi-turn multimodal clinical diagnostic evaluation benchmark to date. ClinMM-Bench contains 1,089 challenging real-world clinical cases and 3,760 medical images across eight specialties. We systematically evaluated 15 representative MLLMs using a two-level evaluation framework that assessed both diagnostic accuracy and diagnostic reasoning quality. Results showed that proprietary models achieved the highest overall diagnostic accuracy, but the proportion of completely correct diagnoses remained limited across all models. In terms of diagnostic reasoning quality, current models can identify plausible diagnostic directi",
  "arxiv_id": "2607.25933",
  "categories": [
    "cs.CL",
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:19:03.000Z"
}
06A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time SeriesQuestion answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to mo{"url":"https://arxiv.org/abs/2607.25947v1","title":"A Cost-Effective Multimodal…
EVENT. cms5r46iID. cms5r46ipagjhkh0cke3jycgxSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25947v1",
  "title": "A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25947",
  "version": "1",
  "abstract": "Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Although recent multimodal time-series large language models (LLMs) have shown considerable promise in general-purpose time-series QA, they remain poorly equipped to model the sparsity, asynchrony, and irregular sampling patterns of clinical observations. To fill this gap, we propose ClinPRISM, a cost-effective multimodal LLM reasoning framework for question answering over ICTS data. First, we devise an irregularity-aware multi-scale encoder to capture sparse clinical evidence at diverse temporal scales. Then, we propose a temporal evidence distiller to integrate representations across these scales and compress them into a small number of LLM-compatible tokens. Moreover, we introduce a progressive alignment strategy that sequentially aligns the irregular trajectories with the LLM's textual embedding space. To facilitate training, we construct 30,000 clinical time series paired with multi-scale descriptions, together with 41,000 instruction-tuning instances spanning 11 tasks. Using a 4-billion-parameter LLM backbone, ClinPRISM achieves state-of-the-art ",
  "arxiv_id": "2607.25947",
  "categories": [
    "cs.AI",
    "cs.CL"
  ],
  "published_at": "2026-07-28T16:33:41.000Z"
}
07MODUS: Decoder-Only Any-to-Any Modeling of Diverse ModalitiesAny-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using{"url":"https://arxiv.org/abs/2607.25948v1","title":"MODUS: Decoder-Only Any-to-…
EVENT. cms5r45yID. cms5r45yuagjfkh0cgoox11knSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25948v1",
  "title": "MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25948",
  "version": "1",
  "abstract": "Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-m",
  "arxiv_id": "2607.25948",
  "categories": [
    "cs.CV",
    "cs.AI",
    "cs.LG"
  ],
  "published_at": "2026-07-28T16:34:23.000Z"
}
08Polistemics: Evaluating LLMs as Information Mediators in Politics & ElectionsAs LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated t{"url":"https://arxiv.org/abs/2607.25953v1","title":"Polistemics: Evaluating LLM…
EVENT. cms5r45fID. cms5r45f7agjdkh0cnrudbillSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25953v1",
  "title": "Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25953",
  "version": "1",
  "abstract": "As LLMs increasingly mediate the political information citizens rely on, there is still no standardized way to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary informational properties such as clarity, noise, and consistency. Applying the benchmark to three state-of-the-art LLMs on the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down under absent, vague, or contradictory information, while flattening the intensity of political language. These failures are likely driven by party priors, influenced by party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.",
  "arxiv_id": "2607.25953",
  "categories": [
    "cs.CL",
    "cs.CY"
  ],
  "published_at": "2026-07-28T16:40:18.000Z"
}
09Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory AllocationMulti-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast {"url":"https://arxiv.org/abs/2607.25956v1","title":"Large Language Model for Op…
EVENT. cms5r44uID. cms5r44uyagjbkh0cddxuin2vSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25956v1",
  "title": "Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25956",
  "version": "1",
  "abstract": "Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast volatility. We study this issue as instance-wise operations research (OR) formulation selection, where each allocation instance is assigned to a solver-executable formulation from a candidate OR expert library. We propose a solver-guided large language model (LLM) framework for OR formulation selection, in which each OR expert corresponds to a MIP formulation encoding a distinct allocation priority. To train the selector, the framework first constructs balanced expert-conditioned supervised fine-tuning (SFT) records for schema learning, and then uses MIP solver evaluation on historical instances to convert solver-evaluated allocation-quality gaps into margin-weighted identity preference optimization (IPO) preferences and per-instance expert-score metadata for reward lookup during group relative policy optimization (GRPO) to assign rewards to sampled responses. Experiments on multi-wareho",
  "arxiv_id": "2607.25956",
  "categories": [
    "cs.AI",
    "math.OC"
  ],
  "published_at": "2026-07-28T16:41:54.000Z"
}
10Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge GraphsWikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raises a practical question: when these modalities disagree, how can we detect and ex{"url":"https://arxiv.org/abs/2607.25959v1","title":"Detecting Knowledge Inconsi…
EVENT. cms5r44bID. cms5r44b8agj9kh0cm5z0ge49SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25959v1",
  "title": "Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25959",
  "version": "1",
  "abstract": "Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is deeply connected but scattered across text, tables, and knowledge graphs. This raises a practical question: when these modalities disagree, how can we detect and explain the conflict? We study this problem as \\emph{modality-level inconsistency detection}. We first introduce a taxonomy of cross-modal knowledge inconsistencies, covering information granularity differences, direct conflicts, temporal changes, and KG incompleteness. We then present \\textsc{Kontrast}, an automatic framework that uses Text-to-SPARQL and LLM reasoning to compare table-based answers with KG evidence and categorize the resulting inconsistencies. Experiments on various Table-QA datasets show that cross-modal inconsistencies are common and informative. They reveal not only true knowledge conflicts, but also missing KG structure and temporal mismatches while being limited by Text-to-SPARQL errors and noise. Our analysis shows that text, tables, and KGs can complement and correct one another through systematic comparison. \\textsc{Kontrast} provides a practical tool for large-sc",
  "arxiv_id": "2607.25959",
  "categories": [
    "cs.CL",
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:43:56.000Z"
}
showing 1–10 of 929older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above