New Datasets
topic · knowledge/datasets-published
§01
about
Top-downloaded datasets on the Hugging Face Hub.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 143 events in this window (179 total on topic). adjust the range or clear it with ALL.
range
01allenai/c4 · 1,540,762 downloadsC4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We pre{"sha":"1588ec454efa1a09f29cd18ddd04fe05fc8653a2","tags":["task_categories:text-…
EVENT. cmrsii36ID. cmrsii36s6zs9kh0cimx4qacmSRC. key:cmpxakb6…
{
"sha": "1588ec454efa1a09f29cd18ddd04fe05fc8653a2",
"tags": [
"task_categories:text-generation",
"task_categories:fill-mask",
"task_ids:language-modeling",
"task_ids:masked-language-modeling",
"annotations_creators:no-annotation",
"language_creators:found",
"multilinguality:multilingual",
"source_datasets:original",
"language:af",
"language:am",
"language:ar",
"language:az",
"language:be",
"language:bg",
"language:bn",
"language:ca",
"language:ceb",
"language:co",
"language:cs",
"language:cy",
"language:da",
"language:de",
"language:el",
"language:en",
"language:eo",
"language:es",
"language:et",
"language:eu",
"language:fa",
"language:fi"
],
"gated": false,
"likes": 618,
"license": "odc-by",
"private": false,
"language": [
"af",
"am",
"ar",
"az",
"be",
"bg",
"bn",
"ca",
"ceb",
"co",
"cs",
"cy",
"da",
"de",
"el",
"en",
"eo",
"es",
"et",
"eu",
"fa",
"fi",
"fil",
"fr",
"fy",
"ga",
"gd",
"gl",
"gu",
"ha",
"haw",
"he",
"hi",
"hmn",
"ht",
"hu",
"hy",
"id",
"ig",
"is",
"it",
"iw",
"ja",
"jv",
"ka",
"kk",
"km",
"kn",
"ko",
"ku",
"ky",
"la",
"lb",
"lo",
"lt",
"lv",
"mg",
"mi",
"mk",
"ml",
"mn",
"mr",
"ms",
"mt",
"my",
"ne",
"nl",
"no",
"ny",
"pa",
"pl",
"ps",
"pt",
"ro",
"ru",
"sd",
"si",
"sk",
"sl",
"sm",
"sn",
"so",
"sq",
"sr",
"st",
"su",
"sv",
"sw",
"ta",
"te",
"tg",
"th",
"tr",
"uk",
"und",
"ur",
"uz",
"vi",
"xh",
"yi",
"yo",
"zh",
"zu"
],
"downloads": 1540762,
"publisher": "allenai",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/allenai/c4",
"dataset_url": "https://huggingface.co/datasets/allenai/c4",
"description": "C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: \"https://commoncrawl.org\". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": "C4",
"card_license": [
"odc-by"
],
"card_languages": [
"af",
"am",
"ar",
"az",
"be"
],
"size_categories": [
"10B<n<100B"
],
"task_categories": [
"text-generation",
"fill-mask"
],
"last_modified_on": "2024-01-09T19:14:03.000Z",
"observation_week": "2026-W30",
"card_task_categories": [
"text-generation",
"fill-mask"
]
}02hf-doc-build/doc-build-dev · 1,609,482 downloadsThis is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider rep{"sha":"494fce0367580237ed1073a60ac801abb072f521","tags":["license:mit","region:…
EVENT. cmrsii2oID. cmrsii2og6zs7kh0clj0kcqkjSRC. key:cmpxakb6…
{
"sha": "494fce0367580237ed1073a60ac801abb072f521",
"tags": [
"license:mit",
"region:us",
"documentation"
],
"gated": false,
"likes": 46,
"license": "mit",
"private": false,
"language": [],
"downloads": 1609482,
"publisher": "hf-doc-build",
"created_on": "2022-11-08T09:03:37.000Z",
"dataset_id": "huggingface.co/datasets/hf-doc-build/doc-build-dev",
"dataset_url": "https://huggingface.co/datasets/hf-doc-build/doc-build-dev",
"description": "This is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider repo.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": "HF Documentation (PRs)",
"card_license": "mit",
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-04-20T08:59:22.000Z",
"observation_week": "2026-W30",
"card_task_categories": []
}03ksolovev/FineNews · 1,698,538 downloads1,698,538 downloads{"sha":"41f8a91dc8c493d0773bc0ebc5158886f5964924","tags":["region:us"],"gated":f…
EVENT. cmrsii26ID. cmrsii2696zs5kh0ckxtjj3hgSRC. key:cmpxakb6…
{
"sha": "41f8a91dc8c493d0773bc0ebc5158886f5964924",
"tags": [
"region:us"
],
"gated": false,
"likes": 23,
"license": null,
"private": false,
"language": [],
"downloads": 1698538,
"publisher": "ksolovev",
"created_on": "2026-03-21T09:09:42.000Z",
"dataset_id": "huggingface.co/datasets/ksolovev/FineNews",
"dataset_url": "https://huggingface.co/datasets/ksolovev/FineNews",
"description": null,
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-03-23T23:30:10.000Z",
"observation_week": "2026-W30",
"card_task_categories": []
}04KakologArchives/KakologArchives · 1,996,645 downloadsニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA な{"sha":"a5bd273128bb544aa4f01b910621630b851d9257","tags":["task_categories:text-…
EVENT. cmrsii1nID. cmrsii1n06zs3kh0cksbjbi0rSRC. key:cmpxakb6…
{
"sha": "a5bd273128bb544aa4f01b910621630b851d9257",
"tags": [
"task_categories:text-classification",
"language:ja",
"license:mit",
"region:us"
],
"gated": false,
"likes": 69,
"license": "mit",
"private": false,
"language": [
"ja"
],
"downloads": 1996645,
"publisher": "KakologArchives",
"created_on": "2023-05-12T13:31:56.000Z",
"dataset_id": "huggingface.co/datasets/KakologArchives/KakologArchives",
"dataset_url": "https://huggingface.co/datasets/KakologArchives/KakologArchives",
"description": "ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/KakologArchives/KakologArchives.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": "ニコニコ実況 過去ログアーカイブ",
"card_license": "mit",
"card_languages": [
"ja"
],
"size_categories": [],
"task_categories": [
"text-classification"
],
"last_modified_on": "2026-07-20T00:52:30.000Z",
"observation_week": "2026-W30",
"card_task_categories": [
"text-classification"
]
}05huggingface/documentation-images · 2,237,680 downloadsThis dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng{"sha":"d934715f5b8a0d154a5c09d8d9b15041052b1913","tags":["license:cc-by-nc-sa-4…
EVENT. cmrsii14ID. cmrsii14p6zs1kh0c1hu6v4bdSRC. key:cmpxakb6…
{
"sha": "d934715f5b8a0d154a5c09d8d9b15041052b1913",
"tags": [
"license:cc-by-nc-sa-4.0",
"size_categories:n<1K",
"format:imagefolder",
"modality:image",
"library:datasets",
"library:mlcroissant",
"region:us"
],
"gated": false,
"likes": 167,
"license": "cc-by-nc-sa-4.0",
"private": false,
"language": [],
"downloads": 2237680,
"publisher": "huggingface",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/huggingface/documentation-images",
"dataset_url": "https://huggingface.co/datasets/huggingface/documentation-images",
"description": "This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": null,
"card_license": "cc-by-nc-sa-4.0",
"card_languages": [],
"size_categories": [
"n<1K"
],
"task_categories": [],
"last_modified_on": "2026-07-17T15:44:02.000Z",
"observation_week": "2026-W30",
"card_task_categories": []
}06Benjy/typed_digital_signatures · 3,026,623 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cmrsii0mID. cmrsii0m86zrzkh0c6apkax54SRC. key:cmpxakb6…
{
"sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
"tags": [
"task_categories:image-classification",
"task_categories:zero-shot-image-classification",
"task_categories:image-feature-extraction",
"language:en",
"license:mit",
"size_categories:10K<n<100K",
"modality:image",
"region:us",
"digital-signatures",
"synthetic-data",
"image-classification",
"computer-vision",
"google-fonts",
"handwriting"
],
"gated": false,
"likes": 16,
"license": "mit",
"private": false,
"language": [
"en"
],
"downloads": 3026623,
"publisher": "Benjy",
"created_on": "2025-01-13T19:08:10.000Z",
"dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
"dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
"description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size:… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": "Typed Digital Signatures Dataset",
"card_license": "mit",
"card_languages": [
"en"
],
"size_categories": [
"10K<n<100K"
],
"task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
],
"last_modified_on": "2025-02-05T18:44:13.000Z",
"observation_week": "2026-W30",
"card_task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
]
}07codeparrot/github-code · 5,703,821 downloadsThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuer{"sha":"b5661e6b17396364b2bcf8e68977b0d28e1ebd19","tags":["task_categories:text-…
EVENT. cmrsii03ID. cmrsii03u6zrxkh0chk49fe3vSRC. key:cmpxakb6…
{
"sha": "b5661e6b17396364b2bcf8e68977b0d28e1ebd19",
"tags": [
"task_categories:text-generation",
"task_ids:language-modeling",
"language_creators:crowdsourced",
"language_creators:expert-generated",
"multilinguality:multilingual",
"language:code",
"license:other",
"region:us"
],
"gated": false,
"likes": 398,
"license": "other",
"private": false,
"language": [
"code"
],
"downloads": 5703821,
"publisher": "codeparrot",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/codeparrot/github-code",
"dataset_url": "https://huggingface.co/datasets/codeparrot/github-code",
"description": "The GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": "github-code",
"card_license": [
"other"
],
"card_languages": [
"code"
],
"size_categories": [],
"task_categories": [
"text-generation"
],
"last_modified_on": "2022-10-20T15:01:14.000Z",
"observation_week": "2026-W30",
"card_task_categories": [
"text-generation"
]
}08anisoleai/fineweb-tokenized · 7,233,276 downloadsFineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokeniz{"sha":"ef1311f460b42138d7a2d18f51e9cc38cedda089","tags":["task_categories:text-…
EVENT. cmrsihzkID. cmrsihzkm6zrvkh0cj2b0z2dmSRC. key:cmpxakb6…
{
"sha": "ef1311f460b42138d7a2d18f51e9cc38cedda089",
"tags": [
"task_categories:text-generation",
"language:en",
"license:odc-by",
"size_categories:n>1T",
"modality:tabular",
"modality:text",
"arxiv:2406.17557",
"region:us",
"tabular",
"text",
"pre-training"
],
"gated": false,
"likes": 16,
"license": "odc-by",
"private": false,
"language": [
"en"
],
"downloads": 7233276,
"publisher": "anisoleai",
"created_on": "2026-05-26T06:34:59.000Z",
"dataset_id": "huggingface.co/datasets/anisoleai/fineweb-tokenized",
"dataset_url": "https://huggingface.co/datasets/anisoleai/fineweb-tokenized",
"description": "FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.",
"observed_at": "2026-07-20T00:54:17.948Z",
"pretty_name": "FineWeb Tokenized (AnisoleAI)",
"card_license": "odc-by",
"card_languages": [
"en"
],
"size_categories": [
"n>1T"
],
"task_categories": [
"text-generation"
],
"last_modified_on": "2026-05-29T12:35:53.000Z",
"observation_week": "2026-W30",
"card_task_categories": [
"text-generation"
]
}09k9cli/video-vec2wav2-tokenizer · 856,896 downloadsvideo-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recogni{"sha":"97ef11e9bb9ef199c7e015551408609b90db6637","tags":["region:us"],"gated":f…
EVENT. cmrqdcc1ID. cmrqdcc196gcxkh0c84pv0xhaSRC. key:cmpxakb6…
{
"sha": "97ef11e9bb9ef199c7e015551408609b90db6637",
"tags": [
"region:us"
],
"gated": false,
"likes": 8,
"license": null,
"private": false,
"language": [],
"downloads": 856896,
"publisher": "k9cli",
"created_on": "2026-06-19T10:28:38.000Z",
"dataset_id": "huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer",
"dataset_url": "https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer",
"description": "video-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.",
"observed_at": "2026-07-18T12:54:23.702Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-07-18T10:34:57.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}10XDOF/ABC-130k · 709,413 downloadsABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with{"sha":"29136bc9b9e38d320b00ffcddbbe4cd0e3278c58","tags":["task_categories:robot…
EVENT. cmrnifu2ID. cmrnifu2e5p4nkh0c69zktjseSRC. key:cmpxakb6…
{
"sha": "29136bc9b9e38d320b00ffcddbbe4cd0e3278c58",
"tags": [
"task_categories:robotics",
"language:en",
"license:apache-2.0",
"size_categories:n>1T",
"arxiv:2606.27375",
"region:us",
"robotics",
"manipulation",
"imitation-learning",
"bimanual",
"teleoperation",
"mcap"
],
"gated": "auto",
"likes": 87,
"license": "apache-2.0",
"private": false,
"language": [
"en"
],
"downloads": 709413,
"publisher": "XDOF",
"created_on": "2026-06-15T22:33:18.000Z",
"dataset_id": "huggingface.co/datasets/XDOF/ABC-130k",
"dataset_url": "https://huggingface.co/datasets/XDOF/ABC-130k",
"description": "ABC-130k ABC-130k is the largest open-source robot teleoperation dataset. It contains bimanual manipulation trajectories collected on two-arm YAM stations. Episodes are distributed as MCAP files, with subtask annotations kept as separate artifacts so they can be revised or extended independently of the underlying episode data. For details on the accompanying paper, see abc.bot. Please see the GitHub repo here for code to train and deploy with this dataset. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/XDOF/ABC-130k.",
"observed_at": "2026-07-16T12:53:46.616Z",
"pretty_name": "ABC",
"card_license": "apache-2.0",
"card_languages": [
"en"
],
"size_categories": [
"n>1T"
],
"task_categories": [
"robotics"
],
"last_modified_on": "2026-07-02T20:47:44.000Z",
"observation_week": "2026-W29",
"card_task_categories": [
"robotics"
]
}showing 1–10 of 143older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above