New Datasets

topic · knowledge/datasets-published
DOC.
knowledge/datasets-published
REV.
179 evt
DATE.
02-JUN-2026
§01

about

Top-downloaded datasets on the Hugging Face Hub.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing the 10 most recent of 179 total events on this topic. apply a date range to scope the list.

range
iso 8601 utc
iso 8601 utc
01genrobot2025/10Kh-RealOmin-OpenData · 786,355 downloadsBoasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,00{"sha":"fcbc0d38550e134f273426aa7c9cc2b491270bc4","tags":["task_categories:robot…
EVENT. cms7ioamID. cms7ioam8ax7hkh0cyvsah1r0SRC. key:cmpxakb6
{
  "sha": "fcbc0d38550e134f273426aa7c9cc2b491270bc4",
  "tags": [
    "task_categories:robotics",
    "task_categories:reinforcement-learning",
    "language:en",
    "language:zh",
    "license:cc-by-sa-4.0",
    "size_categories:n>1T",
    "modality:video",
    "region:us",
    "agent",
    "robotic",
    "real-world",
    "dual-arm",
    "video",
    "vla",
    "embodied intelligence"
  ],
  "gated": "auto",
  "likes": 238,
  "license": "cc-by-sa-4.0",
  "private": false,
  "language": [
    "en",
    "zh"
  ],
  "downloads": 786355,
  "publisher": "genrobot2025",
  "created_on": "2025-12-31T07:17:19.000Z",
  "dataset_id": "huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "dataset_url": "https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "description": "Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.",
  "observed_at": "2026-07-30T12:55:43.891Z",
  "pretty_name": null,
  "card_license": "cc-by-sa-4.0",
  "card_languages": [
    "en",
    "zh"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics",
    "reinforcement-learning"
  ],
  "last_modified_on": "2026-04-24T05:02:26.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "robotics",
    "reinforcement-learning"
  ]
}
02IPEC-COMMUNITY/language_table_lerobot · 1,120,274 downloadsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "xarm", "total_episodes": 442226, "total_frames": 7045476, "total_tasks": 127605, {"sha":"634ac1e50023777c21b299d95aaf2bfdc8514ae9","tags":["task_categories:robot…
EVENT. cms7ioa2ID. cms7ioa20ax7fkh0cup83dk6gSRC. key:cmpxakb6
{
  "sha": "634ac1e50023777c21b299d95aaf2bfdc8514ae9",
  "tags": [
    "task_categories:robotics",
    "license:apache-2.0",
    "region:us",
    "LeRobot",
    "language_table",
    "rlds",
    "openx",
    "xarm"
  ],
  "gated": false,
  "likes": 11,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 1120274,
  "publisher": "IPEC-COMMUNITY",
  "created_on": "2025-03-10T02:03:26.000Z",
  "dataset_id": "huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot",
  "dataset_url": "https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot",
  "description": "This dataset was created using LeRobot. Dataset Structure meta/info.json: { \"codebase_version\": \"v2.0\", \"robot_type\": \"xarm\", \"total_episodes\": 442226, \"total_frames\": 7045476, \"total_tasks\": 127605, \"total_videos\": 442226, \"total_chunks\": 443, \"chunks_size\": 1000, \"fps\": 10, \"splits\": { \"train\": \"0:442226\" }, \"data_path\": \"data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet\", \"video_path\":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/language_table_lerobot.",
  "observed_at": "2026-07-30T12:55:43.891Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2025-03-20T11:33:45.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "robotics"
  ]
}
03HennyPr/ps2_hf2 · 782,263 downloadsn<1K · 782,263 downloads{"sha":"e831be1a0eeb18dbfd99aac845da6bb4c271d08b","tags":["size_categories:n<1K"…
EVENT. cms2ip31ID. cms2ip31x9m91kh0clkrzc31nSRC. key:cmpxakb6
{
  "sha": "e831be1a0eeb18dbfd99aac845da6bb4c271d08b",
  "tags": [
    "size_categories:n<1K",
    "format:text",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 32,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 782263,
  "publisher": "HennyPr",
  "created_on": "2026-03-10T03:07:26.000Z",
  "dataset_id": "huggingface.co/datasets/HennyPr/ps2_hf2",
  "dataset_url": "https://huggingface.co/datasets/HennyPr/ps2_hf2",
  "description": null,
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-04-05T20:00:04.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": []
}
04artur-muratov/multilingual-speech-commands-15lang · 883,811 downloadsMultilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap{"sha":"e38c175c4581649e9018ed3b8c806d5d57b18b19","tags":["language:en","languag…
EVENT. cms2ip2hID. cms2ip2hv9m8zkh0cprmm99evSRC. key:cmpxakb6
{
  "sha": "e38c175c4581649e9018ed3b8c806d5d57b18b19",
  "tags": [
    "language:en",
    "language:ru",
    "language:kk",
    "language:tt",
    "language:ar",
    "language:tr",
    "language:fr",
    "language:de",
    "language:es",
    "language:it",
    "language:ca",
    "language:fa",
    "language:pl",
    "language:nl",
    "language:rw",
    "license:cc-by-4.0",
    "size_categories:1M<n<10M",
    "format:text",
    "modality:audio",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "arxiv:1804.03209",
    "region:us",
    "speech",
    "audio",
    "keyword-spotting",
    "speech-commands",
    "multilingual",
    "low-resource"
  ],
  "gated": false,
  "likes": 11,
  "license": "cc-by-4.0",
  "private": false,
  "language": [
    "en",
    "ru",
    "kk",
    "tt",
    "ar",
    "tr",
    "fr",
    "de",
    "es",
    "it",
    "ca",
    "fa",
    "pl",
    "nl",
    "rw"
  ],
  "downloads": 883811,
  "publisher": "artur-muratov",
  "created_on": "2025-05-27T13:44:01.000Z",
  "dataset_id": "huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang",
  "dataset_url": "https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang",
  "description": "Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": "Multilingual Speech Commands Dataset (15 Languages, Augmented)",
  "card_license": "cc-by-4.0",
  "card_languages": [
    "en",
    "ru",
    "kk",
    "tt",
    "ar"
  ],
  "size_categories": [
    "1M<n<10M"
  ],
  "task_categories": [],
  "last_modified_on": "2025-05-30T01:54:05.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": []
}
05openai/gsm8k · 943,166 downloadsDataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the tas{"sha":"740312add88f781978c0658806c59bc2815b9866","tags":["benchmark:official","…
EVENT. cms2ip1xID. cms2ip1xu9m8xkh0crqhh93cwSRC. key:cmpxakb6
{
  "sha": "740312add88f781978c0658806c59bc2815b9866",
  "tags": [
    "benchmark:official",
    "benchmark:eval-yaml",
    "task_categories:text-generation",
    "annotations_creators:crowdsourced",
    "language_creators:crowdsourced",
    "multilinguality:monolingual",
    "source_datasets:original",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:pandas",
    "library:polars",
    "library:mlcroissant",
    "arxiv:2110.14168",
    "region:us",
    "math-word-problems"
  ],
  "gated": false,
  "likes": 1466,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 943166,
  "publisher": "openai",
  "created_on": "2022-04-12T10:22:10.000Z",
  "dataset_id": "huggingface.co/datasets/openai/gsm8k",
  "dataset_url": "https://huggingface.co/datasets/openai/gsm8k",
  "description": "Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": "Grade School Math 8K",
  "card_license": [
    "mit"
  ],
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2026-03-23T10:18:13.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "text-generation"
  ]
}
06m-a-p/FineFineWeb · 968,019 downloadsFineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iterati{"sha":"7fd92dc825a75cbff271a5a52eea0eda91a2c112","tags":["task_categories:text-…
EVENT. cms2ip1dID. cms2ip1dq9m8vkh0ciy7xb37vSRC. key:cmpxakb6
{
  "sha": "7fd92dc825a75cbff271a5a52eea0eda91a2c112",
  "tags": [
    "task_categories:text-classification",
    "task_categories:text-generation",
    "language:en",
    "license:apache-2.0",
    "size_categories:1B<n<10B",
    "modality:tabular",
    "modality:text",
    "region:us"
  ],
  "gated": false,
  "likes": 164,
  "license": "apache-2.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 968019,
  "publisher": "m-a-p",
  "created_on": "2024-12-14T12:46:33.000Z",
  "dataset_id": "huggingface.co/datasets/m-a-p/FineFineWeb",
  "dataset_url": "https://huggingface.co/datasets/m-a-p/FineFineWeb",
  "description": "FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "1B<n<10B"
  ],
  "task_categories": [
    "text-classification",
    "text-generation"
  ],
  "last_modified_on": "2024-12-19T11:34:03.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "text-classification",
    "text2text-generation",
    "text-generation"
  ]
}
07xlangai/ubuntu_osworld_file_cache · 1,144,610 downloadsOSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld {"sha":"711e0811642364e7aa8f10a8918367d0b626d578","tags":["license:apache-2.0","…
EVENT. cms2ip0tID. cms2ip0tl9m8tkh0czsdsg2r3SRC. key:cmpxakb6
{
  "sha": "711e0811642364e7aa8f10a8918367d0b626d578",
  "tags": [
    "license:apache-2.0",
    "arxiv:2404.07972",
    "region:us"
  ],
  "gated": false,
  "likes": 42,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 1144610,
  "publisher": "xlangai",
  "created_on": "2025-05-27T15:19:04.000Z",
  "dataset_id": "huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache",
  "dataset_url": "https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache",
  "description": "OSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-05-07T11:32:14.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": []
}
08nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim · 1,146,809 downloadsPhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robo{"sha":"ea7ac0b68f87da62f1e726771bba0fe74300802f","tags":["task_categories:robot…
EVENT. cms2ip09ID. cms2ip09k9m8rkh0cz2kxjn7sSRC. key:cmpxakb6
{
  "sha": "ea7ac0b68f87da62f1e726771bba0fe74300802f",
  "tags": [
    "task_categories:robotics",
    "license:cc-by-4.0",
    "region:us",
    "robotics"
  ],
  "gated": false,
  "likes": 251,
  "license": "cc-by-4.0",
  "private": false,
  "language": [],
  "downloads": 1146809,
  "publisher": "nvidia",
  "created_on": "2025-03-18T13:59:39.000Z",
  "dataset_id": "huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "dataset_url": "https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "description": "PhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Cross-embodied bimanual manipulation: 9k trajectories Dataset Name #trajectories bimanual_panda_gripper.Threading 1000 bimanual_panda_hand.LiftTray 1000 bimanual_panda_gripper.ThreePieceAssembly 1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": "cc-by-4.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2026-03-05T23:36:40.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "robotics"
  ]
}
09banned-historical-archives/banned-historical-archives · 1,286,938 downloads和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw{"sha":"7ee825405a889ac77b0d404fac868f9813da0c83","tags":["size_categories:n<1K"…
EVENT. cms2iozpID. cms2iozpi9m8pkh0cetvu98ngSRC. key:cmpxakb6
{
  "sha": "7ee825405a889ac77b0d404fac868f9813da0c83",
  "tags": [
    "size_categories:n<1K",
    "format:imagefolder",
    "modality:image",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 64,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 1286938,
  "publisher": "banned-historical-archives",
  "created_on": "2023-12-17T14:47:08.000Z",
  "dataset_id": "huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
  "dataset_url": "https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
  "description": "和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2025-10-19T15:21:40.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": []
}
10hallucinations-leaderboard/results · 1,414,370 downloads1,414,370 downloads{"sha":"30f759225f89edcc18cb403b93f27bf4231223f7","tags":["license:apache-2.0","…
EVENT. cms2ioz5ID. cms2ioz5f9m8nkh0c9f31cqyoSRC. key:cmpxakb6
{
  "sha": "30f759225f89edcc18cb403b93f27bf4231223f7",
  "tags": [
    "license:apache-2.0",
    "region:us"
  ],
  "gated": false,
  "likes": 16,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 1414370,
  "publisher": "hallucinations-leaderboard",
  "created_on": "2023-11-21T11:44:46.000Z",
  "dataset_id": "huggingface.co/datasets/hallucinations-leaderboard/results",
  "dataset_url": "https://huggingface.co/datasets/hallucinations-leaderboard/results",
  "description": null,
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2024-10-31T20:32:52.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": []
}
showing 1–10 of 179older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above