arXiv Papers
topic · knowledge/papers-arxiv
§01
about
Daily AI and ML paper submissions from cs.AI, cs.LG, cs.CL.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 489 events in this window (598 total on topic). adjust the range or clear it with ALL.
range
01Recurrent Sinusoidal INRs for Efficient High-Fidelity RepresentationWe study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectr{"url":"https://arxiv.org/abs/2607.21485","tldr":"We study sinusoidal recurrence…
EVENT. cmrylpg1ID. cmrylpg1x8nbvkh0ch1pdlf2wSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21485",
"tldr": "We study sinusoidal recurrence as an iterative mechanism for harmonic spectral enrichment in implicit neural representations (INRs). Our analysis reveals that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectral support. We realize this principle with a shared sinusoidal block that iteratively refines the latent representation. We empirically validate the resulting spectral behavior against feed-forward IN",
"title": "Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation",
"authors": [
"Hyunmin Cho",
"Jaejun Yoo",
"Kyong Hwan Jin"
],
"upvotes": 5,
"arxiv_id": "2607.21485",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://hyeon-cho.github.io/Harmonic-line-Spectrum/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}02Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM TextSpatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by r{"url":"https://arxiv.org/abs/2607.21072","tldr":"Spatial intelligence is essent…
EVENT. cmrylpfcID. cmrylpfce8nbrkh0c9ikqwtdjSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21072",
"tldr": "Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by reporting precise coordinates or discrete textual symbols. Yet existing spatial reasoning benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch for image-gener",
"title": "Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text",
"authors": [
"Xu Wang",
"Kaixiang Yao",
"Miao Pan",
"Xiaohe Zhou",
"Xuanyu Liu",
"Wenqi Zhang",
"Xuhong Zhang"
],
"upvotes": 11,
"arxiv_id": "2607.21072",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6485bd278d14bcd5cdbb7c8d/1bC31hrpClCXzgEN_SwSr.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": "ZJU-OmniAI",
"project_page": "https://zju-omniai.github.io/ProVisE/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}03LLMs Get Lost in Evolving User IntentAs LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversat{"url":"https://arxiv.org/abs/2607.20734","tldr":"As LLMs become more capable, t…
EVENT. cmrylpemID. cmrylpemo8nbpkh0c1wo95a7nSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20734",
"tldr": "As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user inten",
"title": "LLMs Get Lost in Evolving User Intent",
"authors": [
"Jihoon Tack",
"Philippe Laban",
"Jennifer Neville"
],
"upvotes": 7,
"arxiv_id": "2607.20734",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Microsoft",
"project_page": null,
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}04Sample-Efficient Learning from Agent ExperienceReal-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once t{"url":"https://arxiv.org/abs/2607.21051","tldr":"Real-world agent learning is o…
EVENT. cmrylpdwID. cmrylpdw58nbnkh0c0dc2d9buSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.21051",
"tldr": "Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-efficient way for agents to learn from their own interaction histories, but its gains disappear once that experience is removed from the context. Separately, context distillation provides a mechanism for internalizing contextual information into model weights. However, applying it to agents' interacti",
"title": "Sample-Efficient Learning from Agent Experience",
"authors": [
"Chenhui Gou",
"Haoqin Tu",
"Yunhao Fang",
"Jianfei Cai",
"Hamid Rezatofighi"
],
"upvotes": 4,
"arxiv_id": "2607.21051",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}05K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMsLarge language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept tax{"url":"https://arxiv.org/abs/2605.09635","tldr":"Large language models are incr…
EVENT. cmrylpd6ID. cmrylpd618nblkh0cmkt5l8mjSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2605.09635",
"tldr": "Large language models are increasingly used in K-12 education, but existing benchmarks mainly test exam question answering rather than understanding how curriculum knowledge is structured and visually presented. We call this capability curriculum cognition. It covers prerequisite chains, concept taxonomies, experiment-concept links, pedagogical sequencing, and visual grounding. We introduce K12-KGraph, a curriculum-aligned knowledge graph extracted from official People's Education Press textbook",
"title": "K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs",
"authors": [
"Hao Liang",
"Qihan Lin",
"Zhaoyang Han",
"Xiaochen Ma",
"Zhen Hao Wong",
"Meiyi Qiang",
"Linzhuang Sun",
"Wentao Zhang"
],
"upvotes": 16,
"arxiv_id": "2605.09635",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": "Peking University",
"project_page": "https://haolpku.github.io/K12-KGraph-page/",
"published_at": "2026-07-23T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-24T00:00:00.000Z"
}06ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language ModelsContextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manif{"url":"https://arxiv.org/abs/2607.20092","tldr":"Contextual entrainment is the …
EVENT. cmrxvxs5ID. cmrxvxs598gerkh0cy319e6x1SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20092",
"tldr": "Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that context is relevant, true, or even meaningful. Recently, it has been identified and given a mechanistic account in unimodal language models. Whether and how it manifests in vision-language models (VLMs) is, by contrast, largely unexamined, and the field lacks a purpose-built instrument with which to investigate it. We take the position that studying contextual en",
"title": "ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models",
"authors": [
"Karan Goyal",
"Afreen Hossain",
"Debojyoti Das",
"Vishal Bhutani"
],
"upvotes": 0,
"arxiv_id": "2607.20092",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/67740b331bb239908a43bceb/e5ZmMGmWZ89GiXBC55WoE.png",
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}07DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document OperationsAs autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework un{"url":"https://arxiv.org/abs/2607.19865","tldr":"As autonomous agents rapidly e…
EVENT. cmrxj2a5ID. cmrxj2a5d8crhkh0cbyvnj214SRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.19865",
"tldr": "As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace workflows. In this paper, we introduce DocOps, a deterministically verifiable evaluation framework underpinned by a hierarchical taxonomy that deconstructs document operations inspired by real-world practices into atomic dimensions and escalating workflow complexities. Based on DocOps, we systematica",
"title": "DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations",
"authors": [
"Jiazhen Jiang",
"Boxi Cao",
"Lingyong Yan",
"Yaojie Lu",
"Hongyu Lin",
"Shuaiqiang Wang",
"Dawei Yin",
"Xianpei Han",
"Le Sun"
],
"upvotes": 2,
"arxiv_id": "2607.19865",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://docopsbench.github.io",
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}08ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly r{"url":"https://arxiv.org/abs/2607.20417","tldr":"3D Gaussian Splatting (3DGS) a…
EVENT. cmrx66k9ID. cmrx66k9o89b1kh0cbszqcbjjSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20417",
"tldr": "3D Gaussian Splatting (3DGS) achieves high-quality novel-view synthesis by optimizing freely placed primitives in 3D and adaptively densifying them in under-reconstructed regions. However, this scene-adaptive capacity allocation is largely lost in existing feed-forward 3DGS methods, which commonly regress Gaussians at input pixels and lift them along camera rays. Such pixel-aligned formulations make the number and placement of primitives depend on image resolution and input viewpoints rather tha",
"title": "ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion",
"authors": [
"Cho In",
"Jeonghwan Cho",
"Mijin Yoo",
"Gim Hee Lee",
"Seon Joo Kim"
],
"upvotes": 0,
"arxiv_id": "2607.20417",
"media_url": "https://cdn-uploads.huggingface.co/production/uploads/6039478ab3ecf716b1a5fd4d/Pj7qRGt69lRZwoi86FpxI.mp4",
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://join16.github.io/page-atsplat/",
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}09SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple EmbodimentsPractical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with l{"url":"https://arxiv.org/abs/2607.20207","tldr":"Practical robotic grasping in …
EVENT. cmrx66jqID. cmrx66jqf89azkh0c1vrwym8tSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.20207",
"tldr": "Practical robotic grasping in complex scenes requires both 3D spatial reasoning and alignment with task-specific requirements. Vision-language models (VLMs) offer a natural way to specify these requirements using language, but existing approaches either use a VLM to predict the grasp directly with limited spatial awareness, or train the VLM together with the grasping model, which requires significantly more data and compute. These limitations impede performance and have prevented scaling to mult",
"title": "SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments",
"authors": [
"Yang Xu",
"Gurpreet Singh Mukker",
"Raymond Wang",
"Jasper Gerigk",
"Maria Attarian",
"Igor Gilitschenski"
],
"upvotes": 0,
"arxiv_id": "2607.20207",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": "https://uoft-isl.github.io/seeded-grasp/",
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}10G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object DetectionThis work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotati{"url":"https://arxiv.org/abs/2607.19942","tldr":"This work introduces G-MAD, an…
EVENT. cmrx66j7ID. cmrx66j7989axkh0c071vm59lSRC. key:cmpxakb6…
{
"url": "https://arxiv.org/abs/2607.19942",
"tldr": "This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-le",
"title": "G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection",
"authors": [
"Yechan Kim",
"JongHyun Park",
"Dongho Yoon",
"Namhoon Jung",
"Moongu Jeon"
],
"upvotes": 0,
"arxiv_id": "2607.19942",
"media_url": null,
"ai_summary": null,
"ai_keywords": [],
"organization": null,
"project_page": null,
"published_at": "2026-07-22T00:00:00.000Z",
"submitted_on_daily_at": "2026-07-23T00:00:00.000Z"
}showing 1–10 of 489older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/papers-arxiv.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above