arXiv AI Submissions

topic · knowledge/arxiv-ai
DOC.
knowledge/arxiv-ai
REV.
995 evt
DATE.
05-JUN-2026
SCOPE.
custom
§01

about

Newest cs.AI and cs.CL submissions from the arXiv API.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 938 events in this window (995 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01APEX-AccountingWe introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, spl{"url":"https://arxiv.org/abs/2607.27189v1","title":"APEX-Accounting","source":"…
EVENT. cms76jh2ID. cms76jh2latwzkh0c7gnmupf7SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27189v1",
  "title": "APEX-Accounting",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27189",
  "version": "1",
  "abstract": "We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.",
  "arxiv_id": "2607.27189",
  "categories": [
    "cs.CL",
    "cs.AI",
    "cs.HC"
  ],
  "published_at": "2026-07-29T17:56:49.000Z"
}
02Can AI agents conduct open-ended AI research? Early evidence from two case studiesForecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind pe{"url":"https://arxiv.org/abs/2607.27191v1","title":"Can AI agents conduct open-…
EVENT. cms76jgiID. cms76jgiwatwxkh0cdoo4nyliSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27191v1",
  "title": "Can AI agents conduct open-ended AI research? Early evidence from two case studies",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27191",
  "version": "1",
  "abstract": "Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor re",
  "arxiv_id": "2607.27191",
  "categories": [
    "cs.AI",
    "cs.CY",
    "cs.LG"
  ],
  "published_at": "2026-07-29T17:57:19.000Z"
}
03Mental World ModelingWorld models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially {"url":"https://arxiv.org/abs/2607.27201v1","title":"Mental World Modeling","sou…
EVENT. cms76jfyID. cms76jfynatwvkh0cruhyum5zSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27201v1",
  "title": "Mental World Modeling",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27201",
  "version": "1",
  "abstract": "World models enable a predictive substrate for planning and action, yet existing formulations merely answer a physical question: what/where it is, and how will it evolve. Human behavior, however, is driven by hidden mental state (what a person believes, wants, intends, feels, and considers socially permissible), so a model that tracks the physical scene but not what each agent knows and believes about it predicts the wrong action for the right-looking scene. We formulate Mental World Modeling (MWM), a generic theoretical framework that makes mental variables core components of a world model rather than posthoc rationales: MWM aintains a coupled physical-mental world state, renders a target-specific partial observation, and simulates how candidate actions jointly update both components. We instantiate the framework in MENTIS, a training-free and fully inspectable baseline that decomposes the process into state parsing, target-observation generation, action decomposition, coupled physical and mental transition, and branch-level value evaluation. On a manually constructed, quality-controlled dataset of situated decision scenarios spanning text, image, and sounding-video stories, exper",
  "arxiv_id": "2607.27201",
  "categories": [
    "cs.CL"
  ],
  "published_at": "2026-07-29T17:59:39.000Z"
}
04Messier: A High-Resolution Corpus for Cross-Benchmark Agent EvaluationEvaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified {"url":"https://arxiv.org/abs/2607.25891v1","title":"Messier: A High-Resolution …
EVENT. cms5r4coID. cms5r4cobagk7kh0chsqc86hdSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25891v1",
  "title": "Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25891",
  "version": "1",
  "abstract": "Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts focus on narrow settings, remain limited in scale, or require costly reruns, leaving much of the empirical record incomparable. We introduce Messier, a unified corpus of 957,253 records that span 30 benchmarks, 714 agents, 11,891 tasks, and 74,205 verifiers. Messier consolidates public benchmark scores and supplements them with five-agent runs across six underrepresented professional and scientific domains, including a recent legal benchmark. Each record is standardized by model, scaffold, environment, task, verifier, and aggregation rule, with SOC/NAICS classifications for occupational and industry analysis. Using this corpus, we show frontier progress is uneven across benchmark types, with \"function calling\" saturated, \"programming\" improving the fastest, and \"enterprise workflows\" remaining the most challenging. Furthermore, counterfactual rescoring shows that strict all-pass aggregation in multi-verifier tasks can obscure progress and artificially alter agent rankings. From these standardized records, we derive capability scales that align ",
  "arxiv_id": "2607.25891",
  "categories": [
    "cs.AI",
    "cs.DB"
  ],
  "published_at": "2026-07-28T15:50:19.000Z"
}
05Interactive Reward Agent: GUI Task Evaluation via Environment-State VerificationGraphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. Howev{"url":"https://arxiv.org/abs/2607.25904v1","title":"Interactive Reward Agent: G…
EVENT. cms5r4c5ID. cms5r4c55agk5kh0cvt7sz6eoSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25904v1",
  "title": "Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25904",
  "version": "1",
  "abstract": "Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% ",
  "arxiv_id": "2607.25904",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:01:38.000Z"
}
06Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language ModelsActivation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an "evaluation-awareness" latent-linearly{"url":"https://arxiv.org/abs/2607.25907v1","title":"Minimizing Targeted Activat…
EVENT. cms5r4bmID. cms5r4bm2agk3kh0cx9i02nc2SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25907v1",
  "title": "Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25907",
  "version": "1",
  "abstract": "Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a fluent prompt so that a chosen internal latent is driven toward zero, with no inference-time model access. Our target is an \"evaluation-awareness\" latent-linearly readable and steerable in recent work-whose control would threaten the validity of safety evaluations if models behave differently when they detect being tested. Adapting Fluent Dreaming / EPO with a negated feature term (GCG-style token optimization plus a self-cross-entropy fluency regularizer, swept over a fluency weight), we suppress the latent under five target constructions-a CAA direction, a subspace norm, an SAE feature, a single MLP neuron, and a behavioral logit-on Llama-3.2-3B and Llama-3.1-8B. The latent is robustly suppressible ($z\\approx-7$), and a causally-validated Llama Scope SAE feature can be fully and selectively turned off. But our controls tell a cautionary story about the CAA direction: a placebo random direction is suppressed just as hard and shifts behavior just as far, and when we hold a real eval passage in context and optimize only a prefix, suppressing the e",
  "arxiv_id": "2607.25907",
  "categories": [
    "cs.LG",
    "cs.AI",
    "cs.CL"
  ],
  "published_at": "2026-07-28T16:01:48.000Z"
}
07AnnoBench: A Benchmark for Visualization Annotation GenerationAnnotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccura{"url":"https://arxiv.org/abs/2607.25911v1","title":"AnnoBench: A Benchmark for …
EVENT. cms5r4b2ID. cms5r4b2wagk1kh0cw2h3pcneSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25911v1",
  "title": "AnnoBench: A Benchmark for Visualization Annotation Generation",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25911",
  "version": "1",
  "abstract": "Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic, and stylistic constraints. Failure to meet any of these conditions severely undermines the utility of an annotation, rendering it challenging to read, inaccurate, or visually discordant. Despite a growing body of annotation tools and automations, no existing benchmark or evaluation framework tests whether these conditions are met because of their scope and annotation not being the focus of their studies. We introduce AnnoBench, a benchmark for visualization annotation that materializes the inherent challenges of this domain in a structured and testable manner. AnnoBench pairs visualizations from professional data journalism and visualization galleries with annotation tasks, spanning four representation formats, five chart description conditions, and two prompt specification levels. The benchmark is executed via VLM-as-a-judge, using models aligned with manual human assessment. We evaluate the benchmark via four one-factor-at-a-time experiments, exploring the effects of input representation, semantic context, and prompt specificity, and model s",
  "arxiv_id": "2607.25911",
  "categories": [
    "cs.HC",
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:04:39.000Z"
}
08SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action ModelsVision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial inter{"url":"https://arxiv.org/abs/2607.25912v1","title":"SAM3D-Guided Object-Centric…
EVENT. cms5r4aiID. cms5r4ai8agjzkh0c4rpf85kjSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25912v1",
  "title": "SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25912",
  "version": "1",
  "abstract": "Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $π_0$, using SAM3D as a frozen 3D teacher to provide target-object 3D priors during training. Specifically, we localize task-relevant objects with object recognition models, generate corresponding object masks, and use SAM3D to extract dense object-level 3D representations, which are aligned with intermediate visual features of $π_0$. This enables the policy to internalize target-object 3D information while preserving the original RGB-language-to-action inference pipeline without requiring depth, point clouds, masks, SAM3D, or additional 3D modules at test time. Simulation experiments show consistent improvements, achieving 99.1\\% on LIBERO and an average length of 4.11 on CALVIN. Real-world experiments further demonstrate that our method is particularly effective in long-horizon manipulation scenarios ",
  "arxiv_id": "2607.25912",
  "categories": [
    "cs.RO",
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:05:32.000Z"
}
09Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous NetworksAutonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- {"url":"https://arxiv.org/abs/2607.25914v1","title":"Toward Standardized Cross-V…
EVENT. cms5r49uID. cms5r49ueagjtkh0c6fphclj8SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25914v1",
  "title": "Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25914",
  "version": "1",
  "abstract": "Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust visibility. When a tool from Vendor B is compromised, agents from Vendor A continue invoking it -- unaware of the trust degradation -- causing cascading service impact. We present AgentToolMO, a proposed 3GPP NRM information model for agent tool trust management. The model comprises: a formally defined trust state machine with provable graduated enforcement, damped cascade propagation with bounded convergence, cross-vendor trust notifications via existing Management Services (MnS) interfaces, and retroactive impact assessment through NRM dependency graph traversal. Simulation-based evaluation across multi-vendor topologies shows that standardized cross-vendor notifications reduce blast radius from hours-scale undetected propagation to near-real-time containment bounded by MnS notification delivery, with cascade convergence guaranteed in bounded iterations and sub-linear notification scaling across vendor domains. The framework operates within existing 3GPP management infrastructure, l",
  "arxiv_id": "2607.25914",
  "categories": [
    "cs.AI",
    "cs.CR",
    "cs.NI"
  ],
  "published_at": "2026-07-28T16:06:41.000Z"
}
10Penelope: Localized Latent Recurrence for Efficient Structured ReasoningComplex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning {"url":"https://arxiv.org/abs/2607.25915v1","title":"Penelope: Localized Latent …
EVENT. cms5r49aID. cms5r49agagjrkh0c1g0l5fdgSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.25915v1",
  "title": "Penelope: Localized Latent Recurrence for Efficient Structured Reasoning",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.25915",
  "version": "1",
  "abstract": "Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while redu",
  "arxiv_id": "2607.25915",
  "categories": [
    "cs.AI"
  ],
  "published_at": "2026-07-28T16:06:46.000Z"
}
showing 1–10 of 938older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above