arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 543 events in this window (596 total on topic). adjust the range or clear it with ALL.
range
01ShieldstralWe introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary questio{"url":"https://arxiv.org/abs/2607.25857","tldr":"We introduce Shieldstral, a 3B…
EVENT. cms5qzdjID. cms5qzdjzaghhkh0cty0wsf7qSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25857",
"tldr": "We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one t",
"title": "Shieldstral",
"authors": [
"Antonia Calvi",
"Avinash Sooriyarachchi",
"Giada Pistilli",
"Guillaume Lample",
"Maarten Buyl",
"Maximilian Augustin",
"Maximilian Müller",
"Pierre Stock",
"Tom Bewley",
"Wassim Bouaziz",
"Yimu Pan"
],
"upvotes": 4,
"arxiv_id": "2607.25857",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Mistral AI_",
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}02Wonder: Video World Model Done BetterWe present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously ob{"url":"https://arxiv.org/abs/2607.26037","tldr":"We present Wonder, a general-p…
EVENT. cms5qzd1ID. cms5qzd1iaghfkh0c3r098kzzSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.26037",
"tldr": "We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given an image or a conditional video, Wonder constructs a playable world where users can navigate interactively by moving the camera, discovering unseen regions, and revisiting previously observed areas in real time and over a long-term horizon. Achieving this capability requires a system-level co-design of control method, memory mechanism, and training strategy. We introduce a novel cam",
"title": "Wonder: Video World Model Done Better",
"authors": [
"Jiacong Xu",
"Hanwen Jiang",
"Zhixin Shu",
"Kalyan Sunkavalli",
"Vishal M. Patel",
"Yiqun Mei"
],
"upvotes": 6,
"arxiv_id": "2607.26037",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/J4rEn3_UurIFAFXZfsiHd.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "Adobe",
"project_page": "https://wonder-world-model.github.io/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}03OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMsEmerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leavi{"url":"https://arxiv.org/abs/2607.25669","tldr":"Emerging Omni-modal Large Lang…
EVENT. cms5qzchID. cms5qzchyaghdkh0c78kih3gjSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25669",
"tldr": "Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory and inference costs. Existing compression methods mainly focus on selecting important tokens under fixed budgets, leaving the preceding budget-allocation problem underexplored. We show that direct query-to-audio/video similarity is unreliable for inter-modal budget allocation, and that uniform intra-modal budgets can ",
"title": "OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs",
"authors": [
"Haoyang Huang",
"Wenjie Huang",
"Tianqi Xu",
"Hongyaoxing Gu",
"Kang Tan",
"Yikai Fu",
"Yuhao Shen",
"Tianyu Liu",
"Baolin Zhang",
"Jun Zhang",
"Xinyi Hu",
"Jun Dai"
],
"upvotes": 2,
"arxiv_id": "2607.25669",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}04Visual prompt engineering for video modelsIn the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask w{"url":"https://arxiv.org/abs/2607.25537","tldr":"In the age of foundation model…
EVENT. cms5qzbzID. cms5qzbzeaghbkh0cntqeup7fSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25537",
"tldr": "In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential technique for improving language model performance. Since video models are currently becoming foundation models for visual tasks (e.g., visual reasoning), we here ask whether they similarly benefit from visual prompt engineering: automatically modifying the task image to improve model performance. For example, for a visual physics reasoning task (\"Where does the bal",
"title": "Visual prompt engineering for video models",
"authors": [
"Robert Geirhos",
"Yuxuan Li",
"Thaddäus Wiedemer",
"Neha Kalibhat",
"Zi Wang",
"Mani Malek",
"Oyvind Tafjord",
"Kevin Swersky",
"Been Kim",
"Priyank Jaini"
],
"upvotes": 3,
"arxiv_id": "2607.25537",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "GoogleDeepMind",
"project_page": "https://visual-prompt-engineering.github.io/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}05Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent MemoryLong-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-n{"url":"https://arxiv.org/abs/2607.24368","tldr":"Long-term memory systems store…
EVENT. cms5qzbcID. cms5qzbcaagh7kh0cyrkmhf5sSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24368",
"tldr": "Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association b",
"title": "Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory",
"authors": [
"Ruizhe Li",
"Mingxuan Du",
"Benfeng Xu",
"Zhendong Mao"
],
"upvotes": 20,
"arxiv_id": "2607.24368",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/68959bf9950f311270f0e010/rWh_nxHDqk6CSJDqPBMIJ.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "muset.ai",
"project_page": "https://keep-it-inmind.github.io",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}06PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language ModelsWe introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domai{"url":"https://arxiv.org/abs/2607.24957","tldr":"We introduce PerceptionBench, …
EVENT. cms5qzasID. cms5qzasragh5kh0cke5cc43uSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24957",
"tldr": "We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagno",
"title": "PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models",
"authors": [
"Zichao Lin",
"Yifeng Xie",
"Bowen Qu",
"Haiming Wang",
"Jia Li",
"Haoning Wu",
"Yuhao Dong",
"Zuhao Yang",
"Jinguo Zhu",
"Haoyu Lu",
"Zijia Zhao",
"Tongtian Yue"
],
"upvotes": 5,
"arxiv_id": "2607.24957",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Moonshot AI",
"project_page": null,
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}07HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data AloneLearning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a s{"url":"https://arxiv.org/abs/2607.25895","tldr":"Learning deployable manipulati…
EVENT. cms5qza9ID. cms5qza9eagh3kh0cbplya4t6SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25895",
"tldr": "Learning deployable manipulation policies is bottlenecked by the scarcity of data that is both high-fidelity and scalable. Real-robot teleoperation is accurate but costly to scale; robot-free UMI capture scales readily, and current practice uses the resulting data mainly for pre-training, adding a small real-robot \"anchor\" at post-training. We ask whether raising the fidelity of robot-free UMI data, rather than shrinking the real-robot fraction, can remove that anchor. We present HiFi-UMI, a por",
"title": "HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone",
"authors": [
"Simple AI",
"Yuteng Wei",
"Jinming Ma",
"Jiawei Wang",
"Weitao Zhou",
"Yushen Zuo",
"Ke Rui",
"Minglei Li",
"Jinhao Zhang",
"Zhikang Pan",
"Xiang Wang",
"Haoran Jia"
],
"upvotes": 68,
"arxiv_id": "2607.25895",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/64060b49a577649430bf6974/seT5Eiwx9StEGYtDsvzqS.webm",
"ai_summary": null,
"ai_keywords": [],
"organization": "Simple World Lab",
"project_page": "https://cloud.simpleai.tech/simple-world-lab/hifi-umi/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}08A New Role for Relevance: Guiding Corpus Interaction in Agentic SearchRelevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction ({"url":"https://arxiv.org/abs/2607.24223","tldr":"Relevance is a query-dependent…
EVENT. cms5qz9qID. cms5qz9q0agh1kh0cbbufk1j9SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24223",
"tldr": "Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot localize, compose, or verify the evidence required by complex questions. Direct Corpus Interaction (DCI) enables such fine-grained operations through grep-style exploration, but its relevance-agnostic search can expose useful clues late and delay convergence. Recent advances use relevance to narrow ",
"title": "A New Role for Relevance: Guiding Corpus Interaction in Agentic Search",
"authors": [
"Jiangnan Li",
"Yuqing Li",
"Mo Yu",
"Jinchao Zhang",
"Jie Zhou"
],
"upvotes": 62,
"arxiv_id": "2607.24223",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Tencent",
"project_page": "https://qdcassie-li.github.io/RARG/",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}09Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation ModelStandard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal u{"url":"https://arxiv.org/abs/2607.24904","tldr":"Standard vision-language model…
EVENT. cms5qz96ID. cms5qz96oaggzkh0cemijfc2kSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24904",
"tldr": "Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich regions using motion vectors and residual energ",
"title": "Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model",
"authors": [
"Senqiao Yang",
"Kaichen Zhang",
"Zhaoyang Jia",
"Jinghao Guo",
"Yifei Shen",
"Xinjie Zhang",
"Xiaoyi Zhang",
"Haoqing Wang",
"Xiao Li",
"Peng Zhang",
"Xiang An",
"Yin Xie"
],
"upvotes": 10,
"arxiv_id": "2607.24904",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Microsoft",
"project_page": "https://microsoft.github.io/Mage",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}10ReDesign: Recovering Editable Design Structures from Images via Agentic DecompositionRecovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector geometry, colors, grouping, and layer ordering. We present ReDesign,{"url":"https://arxiv.org/abs/2607.25565","tldr":"Recovering an editable design …
EVENT. cms5qz8mID. cms5qz8mjaggxkh0cojvql42nSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25565",
"tldr": "Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector geometry, colors, grouping, and layer ordering. We present ReDesign, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across modalities. To keep this long decision process reliable despite imperfect tool outputs,",
"title": "ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition",
"authors": [
"Jooyeol Yun",
"Jintae Park",
"Hyesu Lim",
"Junha Hyung",
"Hyungjin Chung",
"Jaegul Choo"
],
"upvotes": 32,
"arxiv_id": "2607.25565",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6369f693bf21b20c5692937b/DdvUj2HnHJGuCtQbHTDfp.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "KAIST AI",
"project_page": "https://jintae-00.github.io/ReDesign/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}showing 1–10 of 543older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above