arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 561 events in this window (570 total on topic). adjust the range or clear it with ALL.
range
01Can AI agents conduct open-ended AI research? Early evidence from two case studiesForecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind pe{"url":"https://arxiv.org/abs/2607.27191","tldr":"Forecasts of explosive AI prog…
EVENT. cms76e17ID. cms76e17iatuxkh0closi62lsSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.27191",
"tldr": "Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\\&D automation. An agent takes on the central, open-ended ",
"title": "Can AI agents conduct open-ended AI research? Early evidence from two case studies",
"authors": [
"Peter Kirgis",
"Sayash Kapoor",
"Andrew Schwartz",
"Stephan Rabanser",
"David Africa",
"Konstantinos Voudouris",
"Viet Nguyen",
"Toby Pilditch",
"Magda Dubois",
"Harry Coppock",
"Cozmin Ududec",
"Nitya Nadgir"
],
"upvotes": 5,
"arxiv_id": "2607.27191",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://cruxevals.com/",
"published_at": "2026-07-29T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}02OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingLarge language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating L{"url":"https://arxiv.org/abs/2607.27155","tldr":"Large language model (LLM) age…
EVENT. cms76e0oID. cms76e0o5atuvkh0co0hefe8iSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.27155",
"tldr": "Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a pr",
"title": "OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding",
"authors": [
"Jingbo Zhou",
"Yusai Zhao",
"Qi Bao",
"Jingjia Cao",
"Zhenghai Chen",
"Chang Gao",
"Kaiqi Guo",
"Muxin Guo",
"Mingxuan Li",
"Xinjiang Lu",
"Yanru Ma",
"Yixiong Xiao"
],
"upvotes": 6,
"arxiv_id": "2607.27155",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://omegause-officeval.github.io/",
"published_at": "2026-07-29T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}03StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security AgentsStealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents{"url":"https://arxiv.org/abs/2607.26314","tldr":"Stealth, the discipline of ach…
EVENT. cms76e03ID. cms76e03tatutkh0cwnql1vmjSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.26314",
"tldr": "Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents acro",
"title": "StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents",
"authors": [
"Ads Dawson",
"Adrian Wood"
],
"upvotes": 2,
"arxiv_id": "2607.26314",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/660f04586d98d685ae4944ad/rLlQyru2rfz0ADzOcQ6Ta.png",
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://stealthbench.com/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}04CAST: Game Solvers as Turn-Level Teachers for LLM AgentsTraining large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could {"url":"https://arxiv.org/abs/2607.25308","tldr":"Training large language models…
EVENT. cms76dzjID. cms76dzjlaturkh0cutshyimxSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25308",
"tldr": "Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state",
"title": "CAST: Game Solvers as Turn-Level Teachers for LLM Agents",
"authors": [
"Yu Wang",
"Yi-Kai Zhang",
"Wentao Shi",
"Ziang Ye",
"Yuchun Miao",
"Yueqing Sun",
"Qi Gu",
"Xunliang Cai",
"Lan-Zhe Guo",
"Han-Jia Ye",
"Fuli Feng"
],
"upvotes": 25,
"arxiv_id": "2607.25308",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "LongCat",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}05DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text SpaceText-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-en{"url":"https://arxiv.org/abs/2607.25675","tldr":"Text-space optimization adapts…
EVENT. cms76dyzID. cms76dyz7atupkh0c193gtkpaSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25675",
"tldr": "Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so the optimized artifacts remain inspectable and the model can be treated as a black box. However, most existing text-space methods keep evaluation fixed. On open-ended tasks, this can become a bottleneck: once the solver improves on the criteria a rubric measures, omitted dimensions remain invisible to the optimization signal. Simply evolving the rubric is also ",
"title": "DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space",
"authors": [
"Jiangwang Chen",
"Zixin Song",
"Junlin Liu",
"Shuaiyu Zhou",
"Haiyan Wu",
"Haihan Shi",
"Chenxi Zhou",
"Hanqing Li",
"Xiao Yang",
"Da Zhu",
"Guanjun Jiang",
"Hai Wan"
],
"upvotes": 42,
"arxiv_id": "2607.25675",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Qwen Business Unit",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}06SkillRise: Agentic Reinforcement Learning for Cross-Task Skill EvolutionLarge language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines w{"url":"https://arxiv.org/abs/2607.26784","tldr":"Large language model agents of…
EVENT. cms76dygID. cms76dyg1atunkh0ccliqwvo9SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.26784",
"tldr": "Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution. We introduce SkillRise, a unified reinforcement learning framework for learning skills across tasks. SkillRise organizes related",
"title": "SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution",
"authors": [
"Zhiyuan Yao",
"Yuxin Chen",
"Zhengxi Lu",
"Zishan Xu",
"Yueqing Sun",
"Yifu Guo",
"Yuquan Lu",
"Zhengzhou Cai",
"Kangning Zhang",
"Zhuowen Han",
"Zi-Han Wang",
"Ziang Ye"
],
"upvotes": 13,
"arxiv_id": "2607.26784",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-29T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}07CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy OptimizationRubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniforml{"url":"https://arxiv.org/abs/2607.25659","tldr":"Rubric-based reinforcement lea…
EVENT. cms76dxwID. cms76dxwwatulkh0cbrd7nghiSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25659",
"tldr": "Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic ch",
"title": "CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization",
"authors": [
"Bo-Wen Zhang",
"Junwei He",
"Wen Wang",
"Song-Lin Lv",
"Wentao Ma",
"Rongyi Lin",
"Shuhan Zhong",
"Lan-Zhe Guo"
],
"upvotes": 19,
"arxiv_id": "2607.25659",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "ByteDance",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}08CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge AcquisitionReal-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context {"url":"https://arxiv.org/abs/2607.25294","tldr":"Real-world tasks often require…
EVENT. cms76dxdID. cms76dxdpatujkh0cm1yb5ea6SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25294",
"tldr": "Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes",
"title": "CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition",
"authors": [
"Lai Wei",
"Chengqi Li",
"Jiapeng Li",
"Ruina Hu",
"Yue Wang",
"Weiran Huang"
],
"upvotes": 32,
"arxiv_id": "2607.25294",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}09SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident ResponseLarge Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise setti{"url":"https://arxiv.org/abs/2607.26791","tldr":"Large Language Model (LLM) age…
EVENT. cms76dwtID. cms76dwtkatuhkh0cpjzmujycSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.26791",
"tldr": "Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess their security capabilities. However, existing cybersecurity benchmarks focus on pre-compromise settings where agents are placed in a clean and idealized environment before an attack occurs. This leaves the post-compromise setting underexplored. To address this gap, we introduce SecRespond, the first",
"title": "SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response",
"authors": [
"Lehan Wang",
"Boli Chen",
"Ruixue Ding",
"Pengjun Xie",
"Jinwei Huang",
"Zhendong Liu",
"Shuo Wang",
"Tao Lei",
"Xin Ouyang",
"Xiaomeng Li"
],
"upvotes": 0,
"arxiv_id": "2607.26791",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Alibaba-NLP",
"project_page": "https://github.com/Alibaba-NLP/qqr/tree/main/data/secrespond",
"published_at": "2026-07-29T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-30T00:00:00.000Z"
}10Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine DetectionHyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspected before targets are found. This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using {"url":"https://arxiv.org/abs/2607.25310","tldr":"Hyperspectral imaging (HSI) is…
EVENT. cms6tk0nID. cms6tk0n0aqspkh0ch4g5cgwfSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25310",
"tldr": "Hyperspectral imaging (HSI) is useful for material discrimination, but operational mine screening also depends on how many false alarms must be inspected before targets are found. This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using spectral angle mapper (SAM), matched filter (MF), adaptive coherence estimator (ACE), and constrained energy minimization (CEM). We compare a ground-measured SVC signature, a fully informed in-scene c",
"title": "Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection",
"authors": [
"Sagar Lekhak",
"Prasanna Reddy Pulakurthi",
"Emmett J. Ientilucci"
],
"upvotes": 1,
"arxiv_id": "2607.25310",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Rochester Institute of Technology",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}showing 1–10 of 561older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above