arXiv AI Submissions

topic · knowledge/arxiv-ai
DOC.
knowledge/arxiv-ai
REV.
995 evt
DATE.
05-JUN-2026
SCOPE.
custom
§01

about

Newest cs.AI and cs.CL submissions from the arXiv API.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 911 events in this window (995 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task StreamsMemory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over boundary-agnostic and evolving task streams exposes a fundamental stability-plasticity dilemma. External retrieval-based memory can rapidly absorb new ev{"url":"https://arxiv.org/abs/2607.26017v1","title":"UniMem: Complementary Episo…
EVENT. cms5r3ziID. cms5r3zibagirkh0cusf9lvbmSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.26017v1",
  "title": "UniMem: Complementary Episodic-to-Parametric Memory for Boundary-Agnostic Task Streams",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.26017",
  "version": "1",
  "abstract": "Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over boundary-agnostic and evolving task streams exposes a fundamental stability-plasticity dilemma. External retrieval-based memory can rapidly absorb new evidence, but it often fails to internalize recurring execution patterns and incurs inference-time retrieval overhead. Parametric memory enables stable and efficient execution once learned, but typically relies on explicit task boundaries and fixed parameter budgets. Inspired by the human brain, which balances plasticity and stability through complementary episodic storage and gradual consolidation, we propose UniMem, a self-routing framework for autonomous memory management. UniMem uses learnable routing tokens as memory controllers, enabling adaptive coordination between complementary memory pathways: novel or sparse tasks are retained in an episodic buffer for retrieval-augmented execution, while recurring and reliable patterns are consolidated into expandable parametric memory. By decoupling task identification from task execution with routing tokens and parametric memory blocks, UniMe",
  "arxiv_id": "2607.26017",
  "categories": [
    "cs.CL"
  ],
  "published_at": "2026-07-28T17:28:21.000Z"
}
02CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot TransferGraph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, coll{"url":"https://arxiv.org/abs/2607.26023v1","title":"CHARM: A Multimodal Graph F…
EVENT. cms5r3yzID. cms5r3yzhagipkh0cecfen0xvSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.26023v1",
  "title": "CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.26023",
  "version": "1",
  "abstract": "Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world graphs associate nodes with text, images, and other modalities, making multimodal graphs essential for representing complex entities and relations. Moreover, collecting labels and adapting models for every new graph domain is costly and often infeasible, motivating zero-shot transfer. Unfortunately, zero-shot transfer on multimodal graphs remains underexplored. Existing GNN-based graph foundation models typically require downstream adaptation, whereas LLM-based graph methods mainly address unimodal graphs or tasks within a single domain. This setting presents two key challenges. First, models must generalize knowledge from individual modalities while capturing transferable cross-modal relations. Second, without target-domain fine-tuning, node representations remain entangled with domain-specific structures and modality-specific characteristics, obscuring shared concepts in unseen domains. To address these challenges, we propose CHARM, a multimodal graph foundation model with hierarchical context modeling for zero-shot transfer. CHARM replaces iso",
  "arxiv_id": "2607.26023",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-28T17:35:26.000Z"
}
03Falling Behind Drives Unsafe Development in an Idealised AI Race ExperimentTechnological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-cons{"url":"https://arxiv.org/abs/2607.26034v1","title":"Falling Behind Drives Unsaf…
EVENT. cms5r3yfID. cms5r3yfbaginkh0cnwu5h8yzSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.26034v1",
  "title": "Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.26034",
  "version": "1",
  "abstract": "Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\\%, 60\\%, or 90\\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behin",
  "arxiv_id": "2607.26034",
  "categories": [
    "cs.AI",
    "cs.CY",
    "cs.GT",
    "econ.GN"
  ],
  "published_at": "2026-07-28T17:44:03.000Z"
}
04Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for r{"url":"https://arxiv.org/abs/2607.26041v1","title":"Desktop-Delta Bench: Do Com…
EVENT. cms5r3xwID. cms5r3xw4agilkh0cznw5t5dkSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.26041v1",
  "title": "Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.26041",
  "version": "1",
  "abstract": "Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50 task domains. DDB trajectories targets 3 failure dimensions- state verification, source tracking, and context-aware control- through 2 complementary tasks: 463 3-frame temporal-ordering instances, including 105 with a cross-trajectory decoy, and 1,550 before-after pairs labeled from 5 actions + its payload. We evaluate 8 closed and open-source model families across 32 ordering and 16 single-action",
  "arxiv_id": "2607.26041",
  "categories": [
    "cs.AI",
    "cs.CV"
  ],
  "published_at": "2026-07-28T17:49:51.000Z"
}
05$π\mathbf{R}^2$: Reactive Real-time Flow PoliciesGeneralist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \emph{reactivity}. Replanning more often would restore it, but {"url":"https://arxiv.org/abs/2607.26055v1","title":"$π\\mathbf{R}^2$: Reactive …
EVENT. cms5r3xbID. cms5r3xbeagihkh0cwt261kj0SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.26055v1",
  "title": "$π\\mathbf{R}^2$: Reactive Real-time Flow Policies",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.26055",
  "version": "1",
  "abstract": "Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chunks run open-loop, so the policy cannot react to sensory input arriving mid-execution, sacrificing \\emph{reactivity}. Replanning more often would restore it, but the perception-to-action pipeline (a large backbone plus multiple denoising steps) is too slow: this \\emph{latency} forbids frequent replanning and leaves committed actions stale, making such policies ill-suited for dynamic, closed-loop control. We present $π\\mathbf{R}^2$, which makes these policies reactive and real-time while retaining large backbones, expressive multi-modal policies, and multi-action prediction. Built on the per-position noise schedule of diffusion forcing, $π\\mathbf{R}^2$ contributes two ideas. First, it splits conditioning into a fast channel (proprioception, fresh every tick) and an asynchronously updated slow channel (vision-language features), so the policy reacts to proprioception within a chunk while tolerating stale vision. Second, a latency-adaptive flow schedule treats in-flight actions as inpainting conditioning and emits actions in one denoising step per c",
  "arxiv_id": "2607.26055",
  "categories": [
    "cs.RO",
    "cs.AI",
    "cs.LG"
  ],
  "published_at": "2026-07-28T17:59:31.000Z"
}
06Pass the Baton: Trajectory-Relayed On-Policy DistillationOn-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable super{"url":"https://arxiv.org/abs/2607.26057v1","title":"Pass the Baton: Trajectory-…
EVENT. cms5r3wsID. cms5r3ws5agifkh0c5uj1jyn5SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.26057v1",
  "title": "Pass the Baton: Trajectory-Relayed On-Policy Distillation",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.26057",
  "version": "1",
  "abstract": "On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students on eight mathematical reasoning benchmarks, Relay-OPD achieves the best or second-best results on every benchmark, outperforming standard OPD by +5.73% and the strongest baselin",
  "arxiv_id": "2607.26057",
  "categories": [
    "cs.CL",
    "cs.AI"
  ],
  "published_at": "2026-07-28T17:59:46.000Z"
}
07TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge GraphsSecurity Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be truste{"url":"https://arxiv.org/abs/2607.24563v1","title":"TRACE-CTI: Auditable Post-E…
EVENT. cms4bnvzID. cms4bnvz1a331kh0cymaolha4SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24563v1",
  "title": "TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24563",
  "version": "1",
  "abstract": "Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outputs remain fallible and are often stored without the evidence, provenance, and validation history needed to decide whether an individual mapping should be trusted. We present TRACE- CTI, a post-extraction claim-governance framework that preserves run-level Predictions, aggregates them into configuration-level GraphAssertions, materializes setup-deduplicated corroboration as ConsensusAssertions, and exposes only GraphAssertions backed by policy-compliant validation grounds. The framework retains native evidence granularity, complete extraction provenance, versioned trust decisions, and non-destructive revocation history. We evaluate TRACE-CTI on two public CTI corpora comprising 65 reports and 5,303 sentences, using a controlled 2 x 3 matrix of retrievers and generator families, incrementally ingested across six GraphVersions. All setups are incorporated without schema modification; provenance paths remain complete, operational scopes remain disjoint, and every trusted GraphAssertion has an active qualifying validation ground. Cross-generator-fam",
  "arxiv_id": "2607.24563",
  "categories": [
    "cs.AI",
    "cs.CR"
  ],
  "published_at": "2026-07-27T15:33:18.000Z"
}
08DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic HashingSemantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional appr{"url":"https://arxiv.org/abs/2607.24567v1","title":"DSCH-Loss: A Dynamic Semant…
EVENT. cms4bnveID. cms4bnveda32zkh0c7d4by736SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24567v1",
  "title": "DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24567",
  "version": "1",
  "abstract": "Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensional data spaces have gained extensive consideration in recent years. Deep learning-based methods offer better semantic capturing capabilities than traditional approaches relying on manual feature engineering. Moreover, they enable a data-driven approach to semantic hashing across diverse data modalities, yielding high-quality cross-modal hash codes within a shared Hamming space. Previous work investigated the properties of this Hamming space and introduced a loss function based on predefined so-called semantic channels with fixed width and Hamming distances derived from label similarities. However, this formulation also introduced discontinuities into the loss landscape, complicating optimization. Based on these observations, we propose a newly designed loss function, Dynamic Semantic Channel Hashing (DSCH), using dynamically sized and positioned semantic channels in order to avoid loss landscape discontinuities. Furthermore, we endorse the use of tie-aware Mean Average Precision (mAP) as evaluation metric as it addresses the ambiguity in sample r",
  "arxiv_id": "2607.24567",
  "categories": [
    "cs.AI",
    "cs.CV",
    "cs.IR"
  ],
  "published_at": "2026-07-27T15:35:52.000Z"
}
09The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video GroundingLarge-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-o{"url":"https://arxiv.org/abs/2607.24570v1","title":"The Visual Bottleneck: Spar…
EVENT. cms4bnutID. cms4bnutxa32vkh0ckgg99e2aSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24570v1",
  "title": "The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24570",
  "version": "1",
  "abstract": "Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violations occur within each video. Processing every frame is infeasible at scale, so systems are constrained to sparse inputs of 8 to 16 frames per video. Yet state-of-the-art multimodal large language models (MLLMs) are pretrained on dense sequences of hundreds of frames, creating a fundamental mismatch between training and deployment conditions. This mismatch causes severe performance collapse: the Qwen3-VL 8B model drops from 56.0% to 22.3% temporal mIoU when frames are reduced to 16, a 60.2% relative degradation. We present a systematic empirical study of training strategies to close this gap for spatial-temporal video grounding. Our results suggest that visual feature extraction is the dominant bottleneck under sparse-frame inputs. Adapting only the final three ViT layers, 4% of total parameters, achieves 68.8% temporal mIoU and surpasses a zero-shot 8B model using dense inputs by 12.8 points. Language model fine-tuning, by contrast, offers negligible or negative returns. A boundary-aware sampling strategy, Hybrid16, further improves temporal mI",
  "arxiv_id": "2607.24570",
  "categories": [
    "cs.CV",
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:36:21.000Z"
}
10LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in SportsLarge language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesiz{"url":"https://arxiv.org/abs/2607.24573v1","title":"LLM-SoccerArena: Benchmarki…
EVENT. cms4bnuaID. cms4bnuaja32tkh0colgoyw3hSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.24573v1",
  "title": "LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.24573",
  "version": "1",
  "abstract": "Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLM",
  "arxiv_id": "2607.24573",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-27T15:38:29.000Z"
}
showing 1–10 of 911older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above