arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 534 events in this window (596 total on topic). adjust the range or clear it with ALL.
range
01ReDesign: Recovering Editable Design Structures from Images via Agentic DecompositionRecovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector geometry, colors, grouping, and layer ordering. We present ReDesign,{"url":"https://arxiv.org/abs/2607.25565","tldr":"Recovering an editable design …
EVENT. cms5qz8mID. cms5qz8mjaggxkh0cojvql42nSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25565",
"tldr": "Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attributes, such as typography, vector geometry, colors, grouping, and layer ordering. We present ReDesign, an agentic framework that grows an editable layer hierarchy by selecting and composing specialized tools across modalities. To keep this long decision process reliable despite imperfect tool outputs,",
"title": "ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition",
"authors": [
"Jooyeol Yun",
"Jintae Park",
"Hyesu Lim",
"Junha Hyung",
"Hyungjin Chung",
"Jaegul Choo"
],
"upvotes": 32,
"arxiv_id": "2607.25565",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6369f693bf21b20c5692937b/DdvUj2HnHJGuCtQbHTDfp.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "KAIST AI",
"project_page": "https://jintae-00.github.io/ReDesign/",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}02Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label ExpansionWe present a reproducible pipeline for mapping Common Vulnerabilities and Exposures (CVEs) to MITRE ATT&CK Enterprise techniques from free-text vulnerability descriptions. Rather than relying on the CWE->CAPEC->ATT&CK derivation chain, whose table-expansion artifacts we quantify, we train a multi-la{"url":"https://arxiv.org/abs/2607.25572","tldr":"We present a reproducible pipe…
EVENT. cms5qz81ID. cms5qz81faggvkh0cjehbgn9xSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.25572",
"tldr": "We present a reproducible pipeline for mapping Common Vulnerabilities and Exposures (CVEs) to MITRE ATT&CK Enterprise techniques from free-text vulnerability descriptions. Rather than relying on the CWE->CAPEC->ATT&CK derivation chain, whose table-expansion artifacts we quantify, we train a multi-label classifier on a curated gold dataset of 1,207 CVEs from expert MITRE Center for Threat-Informed Defense mappings. The resulting model approximately doubles recall@5 compared with a zero-shot embed",
"title": "Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion",
"authors": [
"Cédric Bonhomme",
"Alexandre Dulaunoy"
],
"upvotes": 2,
"arxiv_id": "2607.25572",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Computer Incident Response Center Luxembourg",
"project_page": "https://github.com/vulnerability-lookup/cve-attack-mapping-paper",
"published_at": "2026-07-28T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-29T00:00:00.000Z"
}03WorldDiT: A Unified Diffusion Architecture for World and Action ModelingMany recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a l{"url":"https://arxiv.org/abs/2607.23909","tldr":"Many recent robot policies pur…
EVENT. cms519h4ID. cms519h42a9p3kh0cxa7z75enSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.23909",
"tldr": "Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates continuous action chunks and predicts normalized RGB patch targets from future camera frames. Across four",
"title": "WorldDiT: A Unified Diffusion Architecture for World and Action Modeling",
"authors": [
"Sen Wang",
"R. Gnana Praveen",
"Bidhan Roy",
"Marcos Villagra"
],
"upvotes": 0,
"arxiv_id": "2607.23909",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Bagel Labs",
"project_page": null,
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}04ID-V2V: Identity-Preserving Video RestylizationIn visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restyliza{"url":"https://arxiv.org/abs/2607.22830","tldr":"In visual storytelling, human …
EVENT. cms4bk2eID. cms4bk2eta30pkh0cm6se9e97SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.22830",
"tldr": "In visual storytelling, human performances are central to creative intent and narrative meaning. However, preserving human identity and performance while enabling flexible visual edits remains challenging for generative video models. We formalize this challenge as identity-preserving video restylization, which propagates scene, lighting, and style changes specified by an edited keyframe across a source video, while preserving facial likeness and performance, including expressions, eye gaze, and ",
"title": "ID-V2V: Identity-Preserving Video Restylization",
"authors": [
"Yuancheng Xu",
"Mingming He",
"Pablo Salamanca",
"Li Ma",
"Yash Kant",
"Emmett Steven",
"Paul Debevec",
"Ning Yu"
],
"upvotes": 2,
"arxiv_id": "2607.22830",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6505e2cdad3134ed7e63c8d0/MZwwo-Th4eHseJysqMVSv.jpeg",
"ai_summary": null,
"ai_keywords": [],
"organization": "Netflix",
"project_page": "https://eyeline-labs.github.io/ID-V2V/",
"published_at": "2026-07-24T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-27T00:00:00.000Z"
}05StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsComputer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different stat{"url":"https://arxiv.org/abs/2607.22798","tldr":"Computer-use agents are usuall…
EVENT. cms4bk1uID. cms4bk1uoa30nkh0c1vh8mmvlSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.22798",
"tldr": "Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying program state, e.g., the files, application backends, and DOM that hold the task data. Different states can produce the same pixels, while code can inspect and modify that state directly. StateAct is a code-first, multi-agent harness built around this distinction. Its main agent works directly with p",
"title": "StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents",
"authors": [
"Yan Yang",
"Xiangru Jian",
"Ziyang Luo",
"Zirui Zhao",
"Yutong Dai",
"Ziji Shi",
"Hanshu Yan",
"Jun Hao Liew",
"Silvio Savarese",
"Junnan Li"
],
"upvotes": 41,
"arxiv_id": "2607.22798",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Salesforce AI Research",
"project_page": null,
"published_at": "2026-07-24T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}06DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style IdentificationDriving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specif{"url":"https://arxiv.org/abs/2607.23822","tldr":"Driving style captures stable,…
EVENT. cms4bk1aID. cms4bk1asa30lkh0cqokcg5gwSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.23822",
"tldr": "Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate because drivers are observed in different vehicles, on different roads, and under different conditions, so models may mistake vehicle- or situation-specific regularities for driver-specific style. We introduce DriveDNA, a large-scale naturalistic dataset and benchmark for personalized driving-style modeling, comprising 4,121 drives from 465 drivers acr",
"title": "DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification",
"authors": [
"Yuhang Wang",
"Lingyao Li",
"Hao Zhou"
],
"upvotes": 0,
"arxiv_id": "2607.23822",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "MOTIF-LAB",
"project_page": null,
"published_at": "2026-07-26T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}07Data Pyramid for Embodied ManipulationMultimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this wo{"url":"https://arxiv.org/abs/2607.24744","tldr":"Multimodal foundation models l…
EVENT. cms4bk0qID. cms4bk0qua30jkh0ckr2199uiSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24744",
"tldr": "Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states and actions. These signals can be provided, to varying degrees, by multiple data sources. In this work, we organize the embodied data ecosystem as a \"pyramid\" spanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-lan",
"title": "Data Pyramid for Embodied Manipulation",
"authors": [
"Yifan Ye",
"Yankai Fu",
"Yaoxu Lv",
"Bohan Hou",
"Jun Cen",
"Lingdong Kong",
"Duo Zheng",
"Tianxing Chen",
"Jiaming Liu",
"Ziang Cao",
"Yunfan Lou",
"Wei Chow"
],
"upvotes": 29,
"arxiv_id": "2607.24744",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/62df78222d89ce551ce0f71d/yPr6P6XzsI1mgHJn29aDk.png",
"ai_summary": null,
"ai_keywords": [],
"organization": "Peking University",
"project_page": "https://jasper-aaa.github.io/embodied-data-pyramid/",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}08ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical UnderstandingMultimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with{"url":"https://arxiv.org/abs/2607.24743","tldr":"Multimodal large language mode…
EVENT. cms4bk06ID. cms4bk06aa30hkh0ctbj33aqlSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24743",
"tldr": "Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical un",
"title": "ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding",
"authors": [
"Hangjie Yuan",
"Yichen Qian",
"Zhiwei Tang",
"Xianzhe Xu",
"Lirong Wu",
"Sicheng Yang",
"Jinwang Wang",
"Pengju Wang",
"Zhitao Zeng",
"Yizeng Han",
"Yan Xing",
"Shengxuan Luo"
],
"upvotes": 3,
"arxiv_id": "2607.24743",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "DAMO Academy",
"project_page": "https://github.com/alibaba-damo-academy/ClinFusion",
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}09OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint GenerationRecent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental struc{"url":"https://arxiv.org/abs/2607.23855","tldr":"Recent generative models are m…
EVENT. cms4bjzmID. cms4bjzmea30fkh0c9wj3lprfSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.23855",
"tldr": "Recent generative models are moving beyond silent video or standalone audio synthesis toward the joint generation of synchronized audio and video. Despite this progress, jointly generating audio and video with fine-grained cross-modal correspondence remains challenging due to their fundamental structural differences. Most existing methods use audio and video VAEs trained separately. As a result, the two latent spaces lack cross-modal alignment, leaving the downstream generative model to learn cr",
"title": "OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation",
"authors": [
"Jun Zhan",
"Chen Yang",
"Yitian Gong",
"Donghua Yu",
"Kuangwei Chen",
"Wenbo Zhang",
"Kexin Huang",
"Qi Luo",
"Zhe Xu",
"Ying Zhu",
"Jin Wang",
"Tengyue Zhang"
],
"upvotes": 13,
"arxiv_id": "2607.23855",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "OpenMOSS",
"project_page": "https://openmoss.ai/OmniVAE.github.io/",
"published_at": "2026-07-26T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}10From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic SearchAgentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced pr{"url":"https://arxiv.org/abs/2607.24280","tldr":"Agentic search enables large l…
EVENT. cms4bjz2ID. cms4bjz2ga30dkh0cvhqnz4jhSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.24280",
"tldr": "Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL) provides only sparse supervision. Knowledge distillation can supply denser guidance, and advanced proprietary models with their strong reasoning capabilities are promising teachers. While distilling from proprietary models can densify this supervisory signal, conventional logit-matching is precluded",
"title": "From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search",
"authors": [
"Junlin Liu",
"Jiangwang Chen",
"Zixin Song",
"Shuaiyu Zhou",
"Chunji Lv",
"Hank Wu",
"Kailin Jiang",
"Jinyang Wu",
"Bohan Yu",
"Chenxi Zhou"
],
"upvotes": 42,
"arxiv_id": "2607.24280",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-27T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-28T00:00:00.000Z"
}showing 1–10 of 534older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above