New Datasets

topic · knowledge/datasets-published
DOC.
knowledge/datasets-published
REV.
179 evt
DATE.
02-JUN-2026
SCOPE.
custom
§01

about

Top-downloaded datasets on the Hugging Face Hub.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 134 events in this window (179 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01XDOF/ABC-130k · 709,413 downloadsABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with{"sha":"29136bc9b9e38d320b00ffcddbbe4cd0e3278c58","tags":["task_categories:robot…
EVENT. cmrnifu2ID. cmrnifu2e5p4nkh0c69zktjseSRC. key:cmpxakb6
{
  "sha": "29136bc9b9e38d320b00ffcddbbe4cd0e3278c58",
  "tags": [
    "task_categories:robotics",
    "language:en",
    "license:apache-2.0",
    "size_categories:n>1T",
    "arxiv:2606.27375",
    "region:us",
    "robotics",
    "manipulation",
    "imitation-learning",
    "bimanual",
    "teleoperation",
    "mcap"
  ],
  "gated": "auto",
  "likes": 87,
  "license": "apache-2.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 709413,
  "publisher": "XDOF",
  "created_on": "2026-06-15T22:33:18.000Z",
  "dataset_id": "huggingface.co/datasets/XDOF/ABC-130k",
  "dataset_url": "https://huggingface.co/datasets/XDOF/ABC-130k",
  "description": "ABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with subtask annotations kept as separate artifacts so they can be revised or extended independently of the underlying episode data. For details on the accompanying paper, see abc.bot. Please see the GitHub repo here for code to train and deploy with this dataset. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/XDOF/ABC-130k.",
  "observed_at": "2026-07-16T12:53:46.616Z",
  "pretty_name": "ABC",
  "card_license": "apache-2.0",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2026-07-02T20:47:44.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "robotics"
  ]
}
02nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim · 690,368 downloadsPhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robo{"sha":"ea7ac0b68f87da62f1e726771bba0fe74300802f","tags":["task_categories:robot…
EVENT. cmriihm5ID. cmriihm5o4cufkh0cu0mxpe1nSRC. key:cmpxakb6
{
  "sha": "ea7ac0b68f87da62f1e726771bba0fe74300802f",
  "tags": [
    "task_categories:robotics",
    "license:cc-by-4.0",
    "region:us",
    "robotics"
  ],
  "gated": false,
  "likes": 243,
  "license": "cc-by-4.0",
  "private": false,
  "language": [],
  "downloads": 690368,
  "publisher": "nvidia",
  "created_on": "2025-03-18T13:59:39.000Z",
  "dataset_id": "huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "dataset_url": "https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "description": "PhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Cross-embodied bimanual manipulation: 9k trajectories Dataset Name #trajectories bimanual_panda_gripper.Threading 1000 bimanual_panda_hand.LiftTray 1000 bimanual_panda_gripper.ThreePieceAssembly 1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": "cc-by-4.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2026-03-05T23:36:40.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "robotics"
  ]
}
03mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M · 746,747 downloads🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀 Upload Status All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics 📜 Cite If you find LLaVA-One-V{"sha":"c5218cad785eba7d218137e8ce4997bda568a050","tags":["license:apache-2.0","…
EVENT. cmriihlmID. cmriihlm54cudkh0c26tw38ziSRC. key:cmpxakb6
{
  "sha": "c5218cad785eba7d218137e8ce4997bda568a050",
  "tags": [
    "license:apache-2.0",
    "size_categories:10M<n<100M",
    "format:parquet",
    "modality:image",
    "modality:text",
    "library:datasets",
    "library:dask",
    "library:mlcroissant",
    "library:polars",
    "arxiv:2509.23661",
    "region:us"
  ],
  "gated": false,
  "likes": 78,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 746747,
  "publisher": "mvp-lab",
  "created_on": "2025-09-14T14:42:33.000Z",
  "dataset_id": "huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M",
  "dataset_url": "https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M",
  "description": "🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀 Upload Status All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics 📜 Cite If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers: @misc{an2025llavaonevision15fullyopenframework, title={LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training}… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [
    "10M<n<100M"
  ],
  "task_categories": [],
  "last_modified_on": "2025-11-24T06:32:02.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": []
}
04HennyPr/ps2_hf2 · 770,524 downloadsn<1K · 770,524 downloads{"sha":"e831be1a0eeb18dbfd99aac845da6bb4c271d08b","tags":["size_categories:n<1K"…
EVENT. cmriihl2ID. cmriihl294cubkh0ckt9of4wbSRC. key:cmpxakb6
{
  "sha": "e831be1a0eeb18dbfd99aac845da6bb4c271d08b",
  "tags": [
    "size_categories:n<1K",
    "format:text",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 24,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 770524,
  "publisher": "HennyPr",
  "created_on": "2026-03-10T03:07:26.000Z",
  "dataset_id": "huggingface.co/datasets/HennyPr/ps2_hf2",
  "dataset_url": "https://huggingface.co/datasets/HennyPr/ps2_hf2",
  "description": null,
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-04-05T20:00:04.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": []
}
05openai/gsm8k · 970,622 downloadsDataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the tas{"sha":"740312add88f781978c0658806c59bc2815b9866","tags":["benchmark:official","…
EVENT. cmriihkiID. cmriihki24cu7kh0cs3q4kevjSRC. key:cmpxakb6
{
  "sha": "740312add88f781978c0658806c59bc2815b9866",
  "tags": [
    "benchmark:official",
    "benchmark:eval-yaml",
    "task_categories:text-generation",
    "annotations_creators:crowdsourced",
    "language_creators:crowdsourced",
    "multilinguality:monolingual",
    "source_datasets:original",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:pandas",
    "library:polars",
    "library:mlcroissant",
    "arxiv:2110.14168",
    "region:us",
    "math-word-problems"
  ],
  "gated": false,
  "likes": 1440,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 970622,
  "publisher": "openai",
  "created_on": "2022-04-12T10:22:10.000Z",
  "dataset_id": "huggingface.co/datasets/openai/gsm8k",
  "dataset_url": "https://huggingface.co/datasets/openai/gsm8k",
  "description": "Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": "Grade School Math 8K",
  "card_license": [
    "mit"
  ],
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2026-03-23T10:18:13.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "text-generation"
  ]
}
06xlangai/ubuntu_osworld_file_cache · 1,219,760 downloadsOSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld {"sha":"711e0811642364e7aa8f10a8918367d0b626d578","tags":["license:apache-2.0","…
EVENT. cmriihjxID. cmriihjx84cu3kh0cbv93am7iSRC. key:cmpxakb6
{
  "sha": "711e0811642364e7aa8f10a8918367d0b626d578",
  "tags": [
    "license:apache-2.0",
    "arxiv:2404.07972",
    "region:us"
  ],
  "gated": false,
  "likes": 37,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 1219760,
  "publisher": "xlangai",
  "created_on": "2025-05-27T15:19:04.000Z",
  "dataset_id": "huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache",
  "dataset_url": "https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache",
  "description": "OSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-05-07T11:32:14.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": []
}
07m-a-p/FineFineWeb · 1,270,270 downloadsFineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iterati{"sha":"7fd92dc825a75cbff271a5a52eea0eda91a2c112","tags":["task_categories:text-…
EVENT. cmriihjeID. cmriihje64cu1kh0ctcanplgxSRC. key:cmpxakb6
{
  "sha": "7fd92dc825a75cbff271a5a52eea0eda91a2c112",
  "tags": [
    "task_categories:text-classification",
    "task_categories:text-generation",
    "language:en",
    "license:apache-2.0",
    "size_categories:1B<n<10B",
    "modality:tabular",
    "modality:text",
    "region:us"
  ],
  "gated": false,
  "likes": 156,
  "license": "apache-2.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 1270270,
  "publisher": "m-a-p",
  "created_on": "2024-12-14T12:46:33.000Z",
  "dataset_id": "huggingface.co/datasets/m-a-p/FineFineWeb",
  "dataset_url": "https://huggingface.co/datasets/m-a-p/FineFineWeb",
  "description": "FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "1B<n<10B"
  ],
  "task_categories": [
    "text-classification",
    "text-generation"
  ],
  "last_modified_on": "2024-12-19T11:34:03.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "text-classification",
    "text2text-generation",
    "text-generation"
  ]
}
08genrobot2025/10Kh-RealOmin-OpenData · 1,272,600 downloadsBoasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,00{"sha":"fcbc0d38550e134f273426aa7c9cc2b491270bc4","tags":["task_categories:robot…
EVENT. cmriihivID. cmriihivd4ctzkh0c16pjdendSRC. key:cmpxakb6
{
  "sha": "fcbc0d38550e134f273426aa7c9cc2b491270bc4",
  "tags": [
    "task_categories:robotics",
    "task_categories:reinforcement-learning",
    "language:en",
    "language:zh",
    "license:cc-by-sa-4.0",
    "size_categories:n>1T",
    "modality:video",
    "region:us",
    "agent",
    "robotic",
    "real-world",
    "dual-arm",
    "video",
    "vla",
    "embodied intelligence"
  ],
  "gated": "auto",
  "likes": 225,
  "license": "cc-by-sa-4.0",
  "private": false,
  "language": [
    "en",
    "zh"
  ],
  "downloads": 1272600,
  "publisher": "genrobot2025",
  "created_on": "2025-12-31T07:17:19.000Z",
  "dataset_id": "huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "dataset_url": "https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "description": "Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": "cc-by-sa-4.0",
  "card_languages": [
    "en",
    "zh"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics",
    "reinforcement-learning"
  ],
  "last_modified_on": "2026-04-24T05:02:26.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "robotics",
    "reinforcement-learning"
  ]
}
09banned-historical-archives/banned-historical-archives · 1,343,237 downloads和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw{"sha":"7ee825405a889ac77b0d404fac868f9813da0c83","tags":["size_categories:n<1K"…
EVENT. cmriihicID. cmriihicl4ctxkh0cbvt1na6xSRC. key:cmpxakb6
{
  "sha": "7ee825405a889ac77b0d404fac868f9813da0c83",
  "tags": [
    "size_categories:n<1K",
    "format:imagefolder",
    "modality:image",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 55,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 1343237,
  "publisher": "banned-historical-archives",
  "created_on": "2023-12-17T14:47:08.000Z",
  "dataset_id": "huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
  "dataset_url": "https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
  "description": "和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2025-10-19T15:21:40.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": []
}
10Salesforce/wikitext · 1,358,255 downloadsDataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia{"sha":"b08601e04326c79dfdd32d625aee71d232d685c3","tags":["task_categories:text-…
EVENT. cmriihhtID. cmriihhtu4ctvkh0clfe3jr4fSRC. key:cmpxakb6
{
  "sha": "b08601e04326c79dfdd32d625aee71d232d685c3",
  "tags": [
    "task_categories:text-generation",
    "task_categories:fill-mask",
    "task_ids:language-modeling",
    "task_ids:masked-language-modeling",
    "annotations_creators:no-annotation",
    "language_creators:crowdsourced",
    "multilinguality:monolingual",
    "source_datasets:original",
    "language:en",
    "license:cc-by-sa-3.0",
    "license:gfdl",
    "size_categories:1M<n<10M",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:dask",
    "library:polars",
    "library:mlcroissant",
    "arxiv:1609.07843",
    "region:us"
  ],
  "gated": false,
  "likes": 741,
  "license": "cc-by-sa-3.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 1358255,
  "publisher": "Salesforce",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/Salesforce/wikitext",
  "dataset_url": "https://huggingface.co/datasets/Salesforce/wikitext",
  "description": "Dataset Card for \"wikitext\" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": "WikiText",
  "card_license": [
    "cc-by-sa-3.0",
    "gfdl"
  ],
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "1M<n<10M"
  ],
  "task_categories": [
    "text-generation",
    "fill-mask"
  ],
  "last_modified_on": "2024-01-04T16:49:18.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "text-generation",
    "fill-mask"
  ]
}
showing 1–10 of 134older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above