New Datasets

topic · knowledge/datasets-published
DOC.
knowledge/datasets-published
REV.
179 evt
DATE.
02-JUN-2026
SCOPE.
custom
§01

about

Top-downloaded datasets on the Hugging Face Hub.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 161 events in this window (179 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01huggingface/documentation-images · 1,925,074 downloadsThis dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng{"sha":"ae19ef9f13a7f4796c96891933fb0b4ed1ddd25b","tags":["license:cc-by-nc-sa-4…
EVENT. cms2iou0ID. cms2iou0e9m83kh0co6rvjcnnSRC. key:cmpxakb6
{
  "sha": "ae19ef9f13a7f4796c96891933fb0b4ed1ddd25b",
  "tags": [
    "license:cc-by-nc-sa-4.0",
    "size_categories:n<1K",
    "format:imagefolder",
    "modality:image",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 172,
  "license": "cc-by-nc-sa-4.0",
  "private": false,
  "language": [],
  "downloads": 1925074,
  "publisher": "huggingface",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/huggingface/documentation-images",
  "dataset_url": "https://huggingface.co/datasets/huggingface/documentation-images",
  "description": "This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": null,
  "card_license": "cc-by-nc-sa-4.0",
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-07-24T07:27:23.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": []
}
02Benjy/typed_digital_signatures · 2,454,445 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cms2iotgID. cms2iotg99m81kh0cvcdh3dfoSRC. key:cmpxakb6
{
  "sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
  "tags": [
    "task_categories:image-classification",
    "task_categories:zero-shot-image-classification",
    "task_categories:image-feature-extraction",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "modality:image",
    "region:us",
    "digital-signatures",
    "synthetic-data",
    "image-classification",
    "computer-vision",
    "google-fonts",
    "handwriting"
  ],
  "gated": false,
  "likes": 20,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 2454445,
  "publisher": "Benjy",
  "created_on": "2025-01-13T19:08:10.000Z",
  "dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
  "dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
  "description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size:… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": "Typed Digital Signatures Dataset",
  "card_license": "mit",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "image-classification",
    "zero-shot-image-classification",
    "image-feature-extraction"
  ],
  "last_modified_on": "2025-02-05T18:44:13.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "image-classification",
    "zero-shot-image-classification",
    "image-feature-extraction"
  ]
}
03codeparrot/github-code · 5,601,917 downloadsThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuer{"sha":"b5661e6b17396364b2bcf8e68977b0d28e1ebd19","tags":["task_categories:text-…
EVENT. cms2iosvID. cms2iosvy9m7zkh0c56wenhfvSRC. key:cmpxakb6
{
  "sha": "b5661e6b17396364b2bcf8e68977b0d28e1ebd19",
  "tags": [
    "task_categories:text-generation",
    "task_ids:language-modeling",
    "language_creators:crowdsourced",
    "language_creators:expert-generated",
    "multilinguality:multilingual",
    "language:code",
    "license:other",
    "region:us"
  ],
  "gated": false,
  "likes": 412,
  "license": "other",
  "private": false,
  "language": [
    "code"
  ],
  "downloads": 5601917,
  "publisher": "codeparrot",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/codeparrot/github-code",
  "dataset_url": "https://huggingface.co/datasets/codeparrot/github-code",
  "description": "The GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": "github-code",
  "card_license": [
    "other"
  ],
  "card_languages": [
    "code"
  ],
  "size_categories": [],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2022-10-20T15:01:14.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "text-generation"
  ]
}
04anisoleai/fineweb-tokenized · 8,542,380 downloadsFineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokeniz{"sha":"ef1311f460b42138d7a2d18f51e9cc38cedda089","tags":["task_categories:text-…
EVENT. cms2iosbID. cms2iosb09m7xkh0cjlc5pd1cSRC. key:cmpxakb6
{
  "sha": "ef1311f460b42138d7a2d18f51e9cc38cedda089",
  "tags": [
    "task_categories:text-generation",
    "language:en",
    "license:odc-by",
    "size_categories:n>1T",
    "format:parquet",
    "modality:tabular",
    "modality:text",
    "library:datasets",
    "library:dask",
    "library:polars",
    "library:mlcroissant",
    "arxiv:2406.17557",
    "region:us",
    "tabular",
    "text",
    "pre-training"
  ],
  "gated": false,
  "likes": 21,
  "license": "odc-by",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 8542380,
  "publisher": "anisoleai",
  "created_on": "2026-05-26T06:34:59.000Z",
  "dataset_id": "huggingface.co/datasets/anisoleai/fineweb-tokenized",
  "dataset_url": "https://huggingface.co/datasets/anisoleai/fineweb-tokenized",
  "description": "FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.",
  "observed_at": "2026-07-27T00:57:16.793Z",
  "pretty_name": "FineWeb Tokenized (AnisoleAI)",
  "card_license": "odc-by",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2026-05-29T12:35:53.000Z",
  "observation_week": "2026-W31",
  "card_task_categories": [
    "text-generation"
  ]
}
05artur-muratov/multilingual-speech-commands-15lang · 864,475 downloadsMultilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap{"sha":"e38c175c4581649e9018ed3b8c806d5d57b18b19","tags":["language:en","languag…
EVENT. cmryy1mlID. cmryy1mlf8qj1kh0c8xhjgsxwSRC. key:cmpxakb6
{
  "sha": "e38c175c4581649e9018ed3b8c806d5d57b18b19",
  "tags": [
    "language:en",
    "language:ru",
    "language:kk",
    "language:tt",
    "language:ar",
    "language:tr",
    "language:fr",
    "language:de",
    "language:es",
    "language:it",
    "language:ca",
    "language:fa",
    "language:pl",
    "language:nl",
    "language:rw",
    "license:cc-by-4.0",
    "size_categories:1M<n<10M",
    "format:text",
    "modality:audio",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "arxiv:1804.03209",
    "region:us",
    "speech",
    "audio",
    "keyword-spotting",
    "speech-commands",
    "multilingual",
    "low-resource"
  ],
  "gated": false,
  "likes": 10,
  "license": "cc-by-4.0",
  "private": false,
  "language": [
    "en",
    "ru",
    "kk",
    "tt",
    "ar",
    "tr",
    "fr",
    "de",
    "es",
    "it",
    "ca",
    "fa",
    "pl",
    "nl",
    "rw"
  ],
  "downloads": 864475,
  "publisher": "artur-muratov",
  "created_on": "2025-05-27T13:44:01.000Z",
  "dataset_id": "huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang",
  "dataset_url": "https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang",
  "description": "Multilingual Speech Commands Dataset (15 Languages, Augmented) This dataset contains augmented speech command samples in 15 languages, derived from multiple public datasets. Only commands that overlap with the Google Speech Commands (GSC) vocabulary are included, making the dataset suitable for multilingual keyword spotting tasks aligned with GSC-style classification. Audio samples have been augmented using standard audio techniques to improve model robustness (e.g., time-shifting… See the full description on the dataset page: https://huggingface.co/datasets/artur-muratov/multilingual-speech-commands-15lang.",
  "observed_at": "2026-07-24T12:56:05.262Z",
  "pretty_name": "Multilingual Speech Commands Dataset (15 Languages, Augmented)",
  "card_license": "cc-by-4.0",
  "card_languages": [
    "en",
    "ru",
    "kk",
    "tt",
    "ar"
  ],
  "size_categories": [
    "1M<n<10M"
  ],
  "task_categories": [],
  "last_modified_on": "2025-05-30T01:54:05.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
06nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim · 1,115,290 downloadsPhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robo{"sha":"ea7ac0b68f87da62f1e726771bba0fe74300802f","tags":["task_categories:robot…
EVENT. cmrxijwgID. cmrxijwgh8cmtkh0c3867zkxySRC. key:cmpxakb6
{
  "sha": "ea7ac0b68f87da62f1e726771bba0fe74300802f",
  "tags": [
    "task_categories:robotics",
    "license:cc-by-4.0",
    "region:us",
    "robotics"
  ],
  "gated": false,
  "likes": 248,
  "license": "cc-by-4.0",
  "private": false,
  "language": [],
  "downloads": 1115290,
  "publisher": "nvidia",
  "created_on": "2025-03-18T13:59:39.000Z",
  "dataset_id": "huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "dataset_url": "https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "description": "PhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Cross-embodied bimanual manipulation: 9k trajectories Dataset Name #trajectories bimanual_panda_gripper.Threading 1000 bimanual_panda_hand.LiftTray 1000 bimanual_panda_gripper.ThreePieceAssembly 1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.",
  "observed_at": "2026-07-23T12:54:38.061Z",
  "pretty_name": null,
  "card_license": "cc-by-4.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2026-03-05T23:36:40.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "robotics"
  ]
}
07XDOF/ABC-130k · 693,442 downloadsABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with{"sha":"29136bc9b9e38d320b00ffcddbbe4cd0e3278c58","tags":["task_categories:robot…
EVENT. cmrsii9bID. cmrsii9b66zszkh0cahsqwh76SRC. key:cmpxakb6
{
  "sha": "29136bc9b9e38d320b00ffcddbbe4cd0e3278c58",
  "tags": [
    "task_categories:robotics",
    "language:en",
    "license:apache-2.0",
    "size_categories:n>1T",
    "arxiv:2606.27375",
    "region:us",
    "robotics",
    "manipulation",
    "imitation-learning",
    "bimanual",
    "teleoperation",
    "mcap"
  ],
  "gated": "auto",
  "likes": 91,
  "license": "apache-2.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 693442,
  "publisher": "XDOF",
  "created_on": "2026-06-15T22:33:18.000Z",
  "dataset_id": "huggingface.co/datasets/XDOF/ABC-130k",
  "dataset_url": "https://huggingface.co/datasets/XDOF/ABC-130k",
  "description": "ABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with subtask annotations kept as separate artifacts so they can be revised or extended independently of the underlying episode data. For details on the accompanying paper, see abc.bot. Please see the GitHub repo here for code to train and deploy with this dataset. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/XDOF/ABC-130k.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": "ABC",
  "card_license": "apache-2.0",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2026-07-02T20:47:44.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "robotics"
  ]
}
08genrobot2025/10Kh-RealOmin-OpenData · 694,670 downloadsBoasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,00{"sha":"fcbc0d38550e134f273426aa7c9cc2b491270bc4","tags":["task_categories:robot…
EVENT. cmrsii8sID. cmrsii8sn6zsvkh0cpjv45z6pSRC. key:cmpxakb6
{
  "sha": "fcbc0d38550e134f273426aa7c9cc2b491270bc4",
  "tags": [
    "task_categories:robotics",
    "task_categories:reinforcement-learning",
    "language:en",
    "language:zh",
    "license:cc-by-sa-4.0",
    "size_categories:n>1T",
    "modality:video",
    "region:us",
    "agent",
    "robotic",
    "real-world",
    "dual-arm",
    "video",
    "vla",
    "embodied intelligence"
  ],
  "gated": "auto",
  "likes": 231,
  "license": "cc-by-sa-4.0",
  "private": false,
  "language": [
    "en",
    "zh"
  ],
  "downloads": 694670,
  "publisher": "genrobot2025",
  "created_on": "2025-12-31T07:17:19.000Z",
  "dataset_id": "huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "dataset_url": "https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "description": "Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": "cc-by-sa-4.0",
  "card_languages": [
    "en",
    "zh"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics",
    "reinforcement-learning"
  ],
  "last_modified_on": "2026-04-24T05:02:26.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "robotics",
    "reinforcement-learning"
  ]
}
09HennyPr/ps2_hf2 · 739,302 downloadsn<1K · 739,302 downloads{"sha":"e831be1a0eeb18dbfd99aac845da6bb4c271d08b","tags":["size_categories:n<1K"…
EVENT. cmrsii8aID. cmrsii8a96zstkh0cuwfto7yfSRC. key:cmpxakb6
{
  "sha": "e831be1a0eeb18dbfd99aac845da6bb4c271d08b",
  "tags": [
    "size_categories:n<1K",
    "format:text",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 28,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 739302,
  "publisher": "HennyPr",
  "created_on": "2026-03-10T03:07:26.000Z",
  "dataset_id": "huggingface.co/datasets/HennyPr/ps2_hf2",
  "dataset_url": "https://huggingface.co/datasets/HennyPr/ps2_hf2",
  "description": null,
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-04-05T20:00:04.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
10k9cli/video-vec2wav2-tokenizer · 909,987 downloadsvideo-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recogni{"sha":"97ef11e9bb9ef199c7e015551408609b90db6637","tags":["region:us"],"gated":f…
EVENT. cmrsii7rID. cmrsii7ry6zsrkh0cljyn4u0tSRC. key:cmpxakb6
{
  "sha": "97ef11e9bb9ef199c7e015551408609b90db6637",
  "tags": [
    "region:us"
  ],
  "gated": false,
  "likes": 10,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 909987,
  "publisher": "k9cli",
  "created_on": "2026-06-19T10:28:38.000Z",
  "dataset_id": "huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer",
  "dataset_url": "https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer",
  "description": "video-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-07-18T10:34:57.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
showing 1–10 of 161older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above