arXiv AI Submissions

topic · knowledge/arxiv-ai
DOC.
knowledge/arxiv-ai
REV.
995 evt
DATE.
05-JUN-2026
SCOPE.
custom
§01

about

Newest cs.AI and cs.CL submissions from the arXiv API.

§02

recent events

LIVElast event 0s ago30 evt / 1h

showing 10 of 947 events in this window (995 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program SynthesisCoding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolv{"url":"https://arxiv.org/abs/2607.27146v1","title":"MindForge: Teaching Small L…
EVENT. cms76jlpID. cms76jlpjatxjkh0cr8svqronSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27146v1",
  "title": "MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27146",
  "version": "1",
  "abstract": "Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier mo",
  "arxiv_id": "2607.27146",
  "categories": [
    "cs.SE",
    "cs.CL",
    "cs.LG"
  ],
  "published_at": "2026-07-29T17:23:02.000Z"
}
02Anatomy Contextualized Adaption of CT Foundation ModelsCT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual fea{"url":"https://arxiv.org/abs/2607.27154v1","title":"Anatomy Contextualized Adap…
EVENT. cms76jl7ID. cms76jl72atxfkh0c46ueoi64SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27154v1",
  "title": "Anatomy Contextualized Adaption of CT Foundation Models",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27154",
  "version": "1",
  "abstract": "CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals. Fine-grained vision-language pre-training addresses this by aligning anatomy-level visual features with anatomy-specific text, but in doing so discards the global context that whole-volume models provide. Furthermore, existing fine-grained approaches train from scratch, making them computationally expensive. We introduce Anatomy Contextualized Adaptation (ACA), a lightweight framework that adapts frozen CT foundation model representations for anatomy-level vision-language alignment while enhancing global contextualization. ACA uses TotalSegmentator to decompose CT volumes into anatomy-level embeddings, which are refined via a transformer that captures cross-anatomy relationships, and aligned to both per-anatomy and scan-level text extracted from radiology reports. Evaluated on Merlin and CT-RATE, ACA consistently outperforms both the frozen foundation model baselines and existing fine-grained methods in zero-shot finding classification, while requiring less than one hour of trai",
  "arxiv_id": "2607.27154",
  "categories": [
    "cs.CV",
    "cs.AI"
  ],
  "published_at": "2026-07-29T17:32:57.000Z"
}
03OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingLarge language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating L{"url":"https://arxiv.org/abs/2607.27155v1","title":"OmegaUse-OfficeVal: Benchma…
EVENT. cms76jkoID. cms76jkohatxdkh0cqm6nojraSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27155v1",
  "title": "OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27155",
  "version": "1",
  "abstract": "Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, a",
  "arxiv_id": "2607.27155",
  "categories": [
    "cs.AI",
    "cs.CL",
    "cs.HC"
  ],
  "published_at": "2026-07-29T17:33:47.000Z"
}
04SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from ScratchLLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a {"url":"https://arxiv.org/abs/2607.27167v1","title":"SpecFirst: Behavioral Speci…
EVENT. cms76jk6ID. cms76jk67atxbkh0cx1q38ca2SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27167v1",
  "title": "SpecFirst: Behavioral Specification Elicitation as a First-Class Step in Agent-Based Program Synthesis from Scratch",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27167",
  "version": "1",
  "abstract": "LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a behavioral oracle, even frontier models solve fewer than 1% of instances. Existing frameworks conflate documentation reading, behavioral exploration, and code synthesis into a single pass, causing agents to probe insufficiently, lose behavioral intent as context drifts, and propagate early misinterpretations into the final implementation. Inspired by classical requirements engineering, we argue that behavioral specification elicitation should be a first-class phase that precedes implementation. We present SpecFirst, a two-stage framework that forces the specification elicitation before code synthesis. A dedicated spec agent first probes the binary and combines observations with documentation into a structured specification. Next, a code synthesis agent then uses this specification to drive implementation. This decomposition resolves documentation ambiguities before coding begins and prov",
  "arxiv_id": "2607.27167",
  "categories": [
    "cs.SE",
    "cs.CL"
  ],
  "published_at": "2026-07-29T17:42:47.000Z"
}
05Improving Item Discoverability in e-Commerce Search via Related Intent GenerationTraditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of subs{"url":"https://arxiv.org/abs/2607.27172v1","title":"Improving Item Discoverabil…
EVENT. cms76jjnID. cms76jjntatx9kh0cibw1i1ehSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27172v1",
  "title": "Improving Item Discoverability in e-Commerce Search via Related Intent Generation",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27172",
  "version": "1",
  "abstract": "Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-commerce marketplaces and particularly grocery, this paradigm is limiting, as user satisfaction and commercial outcomes depend heavily on the discoverability of substitute, complementary, and thematically related items. In this paper, we present a scalable system for discovery-augmented search that leverages intent-conditioned recall expansion. Our approach generates implicit user intents to expand candidate recall while maintaining relevance. The system addresses the cost-quality tradeoff of generative retrieval through a two-stage hybrid architecture. First, we leverage closed-weight large language models (LLMs) to maximize discoverability for head queries. To extend these benefits to tail queries, we then introduce a finetuned small language model (SLM), trained via LoRA adapters and teacher-student distillation. We evaluate the system using a rigorous dual framework: (a) LLM-as-a-judge metrics validated against human preferences for semantic quality, and (b) end-to-end session-level purchase analysis. Results demonstrate that our approach improv",
  "arxiv_id": "2607.27172",
  "categories": [
    "cs.IR",
    "cs.AI"
  ],
  "published_at": "2026-07-29T17:46:35.000Z"
}
06Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc TeamworkEffective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, a{"url":"https://arxiv.org/abs/2607.27177v1","title":"Partner Capability Estimati…
EVENT. cms76jj5ID. cms76jj53atx7kh0cvbmhg55qSRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27177v1",
  "title": "Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27177",
  "version": "1",
  "abstract": "Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we extend ad-hoc teamwork into a multi-task setting by re-framing it as a problem of joint planning with decentralised execution under hidden partner capabilities. We introduce CE-CM (Capability Estimation via Contextual Models), an approximate Bayesian method that infers task-invariant capability vectors. By using simulation-based sampling, the agent estimates capabilities and induces a contextual Multi-agent Markov Decision Processes for planning. This approach requires no population pre-training and refines its beliefs online from just a few tasks. To account for human unpredictability, we propose CE-CM-Div, an extension that evaluates capability hypotheses against diverse planner rollouts rat",
  "arxiv_id": "2607.27177",
  "categories": [
    "cs.AI",
    "cs.HC",
    "cs.MA"
  ],
  "published_at": "2026-07-29T17:50:39.000Z"
}
07DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code SearchState-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and cura{"url":"https://arxiv.org/abs/2607.27178v1","title":"DenseOn with the LateOn: Fu…
EVENT. cms76jimID. cms76jim3atx5kh0co32yyi92SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27178v1",
  "title": "DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27178",
  "version": "1",
  "abstract": "State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for training retrieval models and study how English supervision transfers to multilingual retrieval through translate-train. We first reconstruct and curate 665M English contrastive pre-training pairs from 1.4B pairs across 34 public sources and build 1.88M supervised fine-tuning pairs with mined hard negatives. Training yields two 149M-parameter models: DenseOn, a single-vector dense model, and LateOn, a ColBERT-style late-interaction model. They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting new state-of-the-art results for this size class. We then translate the validated English data into eight languages, yielding 2.8B pairs with cross-lingual samples, and train mDenseOn and mLateOn, two 307M-parameter models built on mmBERT-base. Despite sharing their backbone, data, and objectives, their representations behave differently: the dense model is strong on English and translated languages but degrades outside translate-train support, whereas the late-interaction model generalizes better to unseen languages and scri",
  "arxiv_id": "2607.27178",
  "categories": [
    "cs.CL",
    "cs.IR"
  ],
  "published_at": "2026-07-29T17:50:51.000Z"
}
08The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-MakingConversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surve{"url":"https://arxiv.org/abs/2607.27179v1","title":"The Social Cost of an AI Te…
EVENT. cms76ji3ID. cms76ji3gatx3kh0cam8ma5e8SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27179v1",
  "title": "The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27179",
  "version": "1",
  "abstract": "Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human teams of three. Across six GCA dimensions and survey outcomes, we find that the AI teammate was the single most talkative and self-cohesive member of every treatment team, yet its contributions carried the least new information and the lowest density. The presence of AI also reshaped communication amongst humans. In AI-human teams, human teammates showed lower responsivity and social impact toward one another and reported lower levels of belonging and status. Greater AI dominance in the conversation was associated with students feeling less valued as team members. Additionally, this social cost is immediate and present at baseline; it does not emerge over the course of the conversation. Drawi",
  "arxiv_id": "2607.27179",
  "categories": [
    "cs.HC",
    "cs.AI",
    "cs.CY"
  ],
  "published_at": "2026-07-29T17:51:32.000Z"
}
09Pangram 4 Technical ReportWe present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, Pangram 4 exhibits sup{"url":"https://arxiv.org/abs/2607.27183v1","title":"Pangram 4 Technical Report"…
EVENT. cms76jhkID. cms76jhkuatx1kh0c0y085yxySRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27183v1",
  "title": "Pangram 4 Technical Report",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27183",
  "version": "1",
  "abstract": "We present Pangram 4, the latest deep-learning-based AI-text classification model from Pangram Labs. We achieve an AUROC of 0.9916 with a false positive rate of 0.0041% and a false negative rate of 0.3396%. In addition to its increased overall accuracy compared with Pangram 3, Pangram 4 exhibits superior out-of-distribution generalization and robustness to adversarial attacks. Another novel contribution of Pangram 4 is its improved ability to distinguish fine-grained edits and mixed AI-human co-authored text. We demonstrate improvements to both boundary detection tasks and the detection of interleaved AI assistance. Finally, we report metrics on standard AI detection benchmarks showing that Pangram 4 achieves state-of-the-art performance on the AI text detection task across a wide variety of settings and domains.",
  "arxiv_id": "2607.27183",
  "categories": [
    "cs.CL"
  ],
  "published_at": "2026-07-29T17:53:01.000Z"
}
10APEX-AccountingWe introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, spl{"url":"https://arxiv.org/abs/2607.27189v1","title":"APEX-Accounting","source":"…
EVENT. cms76jh2ID. cms76jh2latwzkh0c7gnmupf7SRC. key:cmpxakb6
{
  "url": "https://arxiv.org/abs/2607.27189v1",
  "title": "APEX-Accounting",
  "source": "arxiv",
  "pdf_url": "https://arxiv.org/pdf/2607.27189",
  "version": "1",
  "abstract": "We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request.",
  "arxiv_id": "2607.27189",
  "categories": [
    "cs.CL",
    "cs.AI",
    "cs.HC"
  ],
  "published_at": "2026-07-29T17:56:49.000Z"
}
showing 1–10 of 947older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/arxiv-ai.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above