New Datasets
topic · knowledge/datasets-published
§01
about
Top-downloaded datasets on the Hugging Face Hub.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 125 events in this window (179 total on topic). adjust the range or clear it with ALL.
range
01Salesforce/wikitext · 1,358,255 downloadsDataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia{"sha":"b08601e04326c79dfdd32d625aee71d232d685c3","tags":["task_categories:text-…
EVENT. cmriihhtID. cmriihhtu4ctvkh0clfe3jr4fSRC. key:cmpxakb6…
{
"sha": "b08601e04326c79dfdd32d625aee71d232d685c3",
"tags": [
"task_categories:text-generation",
"task_categories:fill-mask",
"task_ids:language-modeling",
"task_ids:masked-language-modeling",
"annotations_creators:no-annotation",
"language_creators:crowdsourced",
"multilinguality:monolingual",
"source_datasets:original",
"language:en",
"license:cc-by-sa-3.0",
"license:gfdl",
"size_categories:1M<n<10M",
"format:parquet",
"modality:text",
"library:datasets",
"library:dask",
"library:polars",
"library:mlcroissant",
"arxiv:1609.07843",
"region:us"
],
"gated": false,
"likes": 741,
"license": "cc-by-sa-3.0",
"private": false,
"language": [
"en"
],
"downloads": 1358255,
"publisher": "Salesforce",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/Salesforce/wikitext",
"dataset_url": "https://huggingface.co/datasets/Salesforce/wikitext",
"description": "Dataset Card for \"wikitext\" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": "WikiText",
"card_license": [
"cc-by-sa-3.0",
"gfdl"
],
"card_languages": [
"en"
],
"size_categories": [
"1M<n<10M"
],
"task_categories": [
"text-generation",
"fill-mask"
],
"last_modified_on": "2024-01-04T16:49:18.000Z",
"observation_week": "2026-W29",
"card_task_categories": [
"text-generation",
"fill-mask"
]
}02ryanmarten/OpenThoughts-1k-sample · 1,458,458 downloads[!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality {"sha":"a82400884621626d41bef89b7604f8054e7e00e0","tags":["size_categories:1K<n<…
EVENT. cmriihhaID. cmriihhaw4cttkh0chxxwudxiSRC. key:cmpxakb6…
{
"sha": "a82400884621626d41bef89b7604f8054e7e00e0",
"tags": [
"size_categories:1K<n<10K",
"format:parquet",
"modality:text",
"library:datasets",
"library:pandas",
"library:mlcroissant",
"library:polars",
"arxiv:2506.04178",
"region:us"
],
"gated": false,
"likes": 37,
"license": null,
"private": false,
"language": [],
"downloads": 1458458,
"publisher": "ryanmarten",
"created_on": "2025-08-30T23:58:46.000Z",
"dataset_id": "huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample",
"dataset_url": "https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample",
"description": "[!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [
"1K<n<10K"
],
"task_categories": [],
"last_modified_on": "2025-08-31T00:33:15.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}03hf-doc-build/doc-build-dev · 1,462,545 downloadsThis is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider rep{"sha":"494fce0367580237ed1073a60ac801abb072f521","tags":["license:mit","region:…
EVENT. cmriihgsID. cmriihgs04ctrkh0cd4ozytskSRC. key:cmpxakb6…
{
"sha": "494fce0367580237ed1073a60ac801abb072f521",
"tags": [
"license:mit",
"region:us",
"documentation"
],
"gated": false,
"likes": 42,
"license": "mit",
"private": false,
"language": [],
"downloads": 1462545,
"publisher": "hf-doc-build",
"created_on": "2022-11-08T09:03:37.000Z",
"dataset_id": "huggingface.co/datasets/hf-doc-build/doc-build-dev",
"dataset_url": "https://huggingface.co/datasets/hf-doc-build/doc-build-dev",
"description": "This is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider repo.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": "HF Documentation (PRs)",
"card_license": "mit",
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-04-20T08:59:22.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}04ayuo/hd_tmp · 1,477,012 downloads1,477,012 downloads{"sha":"0b7fd955a202c05b39cf9bc816c270fbdb391ea4","tags":["region:us"],"gated":f…
EVENT. cmriihg9ID. cmriihg9c4ctpkh0c4lwcpa3nSRC. key:cmpxakb6…
{
"sha": "0b7fd955a202c05b39cf9bc816c270fbdb391ea4",
"tags": [
"region:us"
],
"gated": false,
"likes": 22,
"license": null,
"private": false,
"language": [],
"downloads": 1477012,
"publisher": "ayuo",
"created_on": "2026-04-01T14:22:50.000Z",
"dataset_id": "huggingface.co/datasets/ayuo/hd_tmp",
"dataset_url": "https://huggingface.co/datasets/ayuo/hd_tmp",
"description": null,
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-07-05T09:46:14.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}05allenai/c4 · 1,502,286 downloadsC4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We pre{"sha":"1588ec454efa1a09f29cd18ddd04fe05fc8653a2","tags":["task_categories:text-…
EVENT. cmriihfqID. cmriihfqh4ctnkh0c90todxlbSRC. key:cmpxakb6…
{
"sha": "1588ec454efa1a09f29cd18ddd04fe05fc8653a2",
"tags": [
"task_categories:text-generation",
"task_categories:fill-mask",
"task_ids:language-modeling",
"task_ids:masked-language-modeling",
"annotations_creators:no-annotation",
"language_creators:found",
"multilinguality:multilingual",
"source_datasets:original",
"language:af",
"language:am",
"language:ar",
"language:az",
"language:be",
"language:bg",
"language:bn",
"language:ca",
"language:ceb",
"language:co",
"language:cs",
"language:cy",
"language:da",
"language:de",
"language:el",
"language:en",
"language:eo",
"language:es",
"language:et",
"language:eu",
"language:fa",
"language:fi"
],
"gated": false,
"likes": 610,
"license": "odc-by",
"private": false,
"language": [
"af",
"am",
"ar",
"az",
"be",
"bg",
"bn",
"ca",
"ceb",
"co",
"cs",
"cy",
"da",
"de",
"el",
"en",
"eo",
"es",
"et",
"eu",
"fa",
"fi",
"fil",
"fr",
"fy",
"ga",
"gd",
"gl",
"gu",
"ha",
"haw",
"he",
"hi",
"hmn",
"ht",
"hu",
"hy",
"id",
"ig",
"is",
"it",
"iw",
"ja",
"jv",
"ka",
"kk",
"km",
"kn",
"ko",
"ku",
"ky",
"la",
"lb",
"lo",
"lt",
"lv",
"mg",
"mi",
"mk",
"ml",
"mn",
"mr",
"ms",
"mt",
"my",
"ne",
"nl",
"no",
"ny",
"pa",
"pl",
"ps",
"pt",
"ro",
"ru",
"sd",
"si",
"sk",
"sl",
"sm",
"sn",
"so",
"sq",
"sr",
"st",
"su",
"sv",
"sw",
"ta",
"te",
"tg",
"th",
"tr",
"uk",
"und",
"ur",
"uz",
"vi",
"xh",
"yi",
"yo",
"zh",
"zu"
],
"downloads": 1502286,
"publisher": "allenai",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/allenai/c4",
"dataset_url": "https://huggingface.co/datasets/allenai/c4",
"description": "C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: \"https://commoncrawl.org\". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": "C4",
"card_license": [
"odc-by"
],
"card_languages": [
"af",
"am",
"ar",
"az",
"be"
],
"size_categories": [
"10B<n<100B"
],
"task_categories": [
"text-generation",
"fill-mask"
],
"last_modified_on": "2024-01-09T19:14:03.000Z",
"observation_week": "2026-W29",
"card_task_categories": [
"text-generation",
"fill-mask"
]
}06hallucinations-leaderboard/results · 1,510,578 downloads1,510,578 downloads{"sha":"30f759225f89edcc18cb403b93f27bf4231223f7","tags":["license:apache-2.0","…
EVENT. cmriihf7ID. cmriihf7c4ctlkh0c6dh8kz40SRC. key:cmpxakb6…
{
"sha": "30f759225f89edcc18cb403b93f27bf4231223f7",
"tags": [
"license:apache-2.0",
"region:us"
],
"gated": false,
"likes": 10,
"license": "apache-2.0",
"private": false,
"language": [],
"downloads": 1510578,
"publisher": "hallucinations-leaderboard",
"created_on": "2023-11-21T11:44:46.000Z",
"dataset_id": "huggingface.co/datasets/hallucinations-leaderboard/results",
"dataset_url": "https://huggingface.co/datasets/hallucinations-leaderboard/results",
"description": null,
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": null,
"card_license": "apache-2.0",
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2024-10-31T20:32:52.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}07ksolovev/FineNews · 1,961,580 downloads1,961,580 downloads{"sha":"41f8a91dc8c493d0773bc0ebc5158886f5964924","tags":["region:us"],"gated":f…
EVENT. cmriiheoID. cmriiheot4cthkh0cwz2atxgmSRC. key:cmpxakb6…
{
"sha": "41f8a91dc8c493d0773bc0ebc5158886f5964924",
"tags": [
"region:us"
],
"gated": false,
"likes": 20,
"license": null,
"private": false,
"language": [],
"downloads": 1961580,
"publisher": "ksolovev",
"created_on": "2026-03-21T09:09:42.000Z",
"dataset_id": "huggingface.co/datasets/ksolovev/FineNews",
"dataset_url": "https://huggingface.co/datasets/ksolovev/FineNews",
"description": null,
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-03-23T23:30:10.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}08KakologArchives/KakologArchives · 2,229,125 downloadsニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA な{"sha":"90c312aa6063317cc7e149acbc17a6883a4027ba","tags":["task_categories:text-…
EVENT. cmriihe5ID. cmriihe5b4ctfkh0c7x5u8b8tSRC. key:cmpxakb6…
{
"sha": "90c312aa6063317cc7e149acbc17a6883a4027ba",
"tags": [
"task_categories:text-classification",
"language:ja",
"license:mit",
"region:us"
],
"gated": false,
"likes": 64,
"license": "mit",
"private": false,
"language": [
"ja"
],
"downloads": 2229125,
"publisher": "KakologArchives",
"created_on": "2023-05-12T13:31:56.000Z",
"dataset_id": "huggingface.co/datasets/KakologArchives/KakologArchives",
"dataset_url": "https://huggingface.co/datasets/KakologArchives/KakologArchives",
"description": "ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/KakologArchives/KakologArchives.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": "ニコニコ実況 過去ログアーカイブ",
"card_license": "mit",
"card_languages": [
"ja"
],
"size_categories": [],
"task_categories": [
"text-classification"
],
"last_modified_on": "2026-07-12T19:47:30.000Z",
"observation_week": "2026-W29",
"card_task_categories": [
"text-classification"
]
}09huggingface/documentation-images · 2,882,214 downloadsThis dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng{"sha":"443c57b8d621bcf7a29bc7a7b8e303ebd2615c1e","tags":["license:cc-by-nc-sa-4…
EVENT. cmriihdmID. cmriihdmh4ctdkh0c56huep1dSRC. key:cmpxakb6…
{
"sha": "443c57b8d621bcf7a29bc7a7b8e303ebd2615c1e",
"tags": [
"license:cc-by-nc-sa-4.0",
"size_categories:n<1K",
"format:imagefolder",
"modality:image",
"library:datasets",
"library:mlcroissant",
"region:us"
],
"gated": false,
"likes": 164,
"license": "cc-by-nc-sa-4.0",
"private": false,
"language": [],
"downloads": 2882214,
"publisher": "huggingface",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/huggingface/documentation-images",
"dataset_url": "https://huggingface.co/datasets/huggingface/documentation-images",
"description": "This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": null,
"card_license": "cc-by-nc-sa-4.0",
"card_languages": [],
"size_categories": [
"n<1K"
],
"task_categories": [],
"last_modified_on": "2026-07-10T07:27:05.000Z",
"observation_week": "2026-W29",
"card_task_categories": []
}10Benjy/typed_digital_signatures · 3,109,722 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cmriihd3ID. cmriihd3l4ctbkh0c2nzjkhuvSRC. key:cmpxakb6…
{
"sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
"tags": [
"task_categories:image-classification",
"task_categories:zero-shot-image-classification",
"task_categories:image-feature-extraction",
"language:en",
"license:mit",
"size_categories:10K<n<100K",
"modality:image",
"region:us",
"digital-signatures",
"synthetic-data",
"image-classification",
"computer-vision",
"google-fonts",
"handwriting"
],
"gated": false,
"likes": 11,
"license": "mit",
"private": false,
"language": [
"en"
],
"downloads": 3109722,
"publisher": "Benjy",
"created_on": "2025-01-13T19:08:10.000Z",
"dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
"dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
"description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size:… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
"observed_at": "2026-07-13T00:56:05.317Z",
"pretty_name": "Typed Digital Signatures Dataset",
"card_license": "mit",
"card_languages": [
"en"
],
"size_categories": [
"10K<n<100K"
],
"task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
],
"last_modified_on": "2025-02-05T18:44:13.000Z",
"observation_week": "2026-W29",
"card_task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
]
}showing 1–10 of 125older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above