New Datasets
topic · knowledge/datasets-published
§01
about
Top-downloaded datasets on the Hugging Face Hub.
§02
recent events
LIVElast event 0s ago0 evt / 1h
showing 10 of 80 events in this window (179 total on topic). adjust the range or clear it with ALL.
range
01ayuo/hd_tmp · 1,442,217 downloads1,442,217 downloads{"sha":"f26536be5fdb05cbff767d5a88787f626ec9ea73","tags":["region:us"],"gated":f…
EVENT. cmr2sokhID. cmr2sokhv051jkh0cf8zb1y98SRC. key:cmpxakb6…
{
"sha": "f26536be5fdb05cbff767d5a88787f626ec9ea73",
"tags": [
"region:us"
],
"gated": false,
"likes": 18,
"license": null,
"private": false,
"language": [],
"downloads": 1442217,
"publisher": "ayuo",
"created_on": "2026-04-01T14:22:50.000Z",
"dataset_id": "huggingface.co/datasets/ayuo/hd_tmp",
"dataset_url": "https://huggingface.co/datasets/ayuo/hd_tmp",
"description": null,
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-06-10T22:01:38.000Z",
"observation_week": "2026-W27",
"card_task_categories": []
}02hallucinations-leaderboard/results · 1,470,013 downloads1,470,013 downloads{"sha":"30f759225f89edcc18cb403b93f27bf4231223f7","tags":["license:apache-2.0","…
EVENT. cmr2sojwID. cmr2sojwu051dkh0c4911u5fkSRC. key:cmpxakb6…
{
"sha": "30f759225f89edcc18cb403b93f27bf4231223f7",
"tags": [
"license:apache-2.0",
"region:us"
],
"gated": false,
"likes": 2,
"license": "apache-2.0",
"private": false,
"language": [],
"downloads": 1470013,
"publisher": "hallucinations-leaderboard",
"created_on": "2023-11-21T11:44:46.000Z",
"dataset_id": "huggingface.co/datasets/hallucinations-leaderboard/results",
"dataset_url": "https://huggingface.co/datasets/hallucinations-leaderboard/results",
"description": null,
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": null,
"card_license": "apache-2.0",
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2024-10-31T20:32:52.000Z",
"observation_week": "2026-W27",
"card_task_categories": []
}03codeparrot/github-code · 1,487,171 downloadsThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuer{"sha":"b5661e6b17396364b2bcf8e68977b0d28e1ebd19","tags":["task_categories:text-…
EVENT. cmr2sojbID. cmr2sojbw0519kh0cx3vvm33jSRC. key:cmpxakb6…
{
"sha": "b5661e6b17396364b2bcf8e68977b0d28e1ebd19",
"tags": [
"task_categories:text-generation",
"task_ids:language-modeling",
"language_creators:crowdsourced",
"language_creators:expert-generated",
"multilinguality:multilingual",
"language:code",
"license:other",
"region:us"
],
"gated": false,
"likes": 367,
"license": "other",
"private": false,
"language": [
"code"
],
"downloads": 1487171,
"publisher": "codeparrot",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/codeparrot/github-code",
"dataset_url": "https://huggingface.co/datasets/codeparrot/github-code",
"description": "The GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.",
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": "github-code",
"card_license": [
"other"
],
"card_languages": [
"code"
],
"size_categories": [],
"task_categories": [
"text-generation"
],
"last_modified_on": "2022-10-20T15:01:14.000Z",
"observation_week": "2026-W27",
"card_task_categories": [
"text-generation"
]
}04banned-historical-archives/banned-historical-archives · 1,513,631 downloads和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw{"sha":"7ee825405a889ac77b0d404fac868f9813da0c83","tags":["size_categories:n<1K"…
EVENT. cmr2soiqID. cmr2soiqv0517kh0czy3duv21SRC. key:cmpxakb6…
{
"sha": "7ee825405a889ac77b0d404fac868f9813da0c83",
"tags": [
"size_categories:n<1K",
"format:imagefolder",
"modality:image",
"library:datasets",
"library:mlcroissant",
"region:us"
],
"gated": false,
"likes": 48,
"license": null,
"private": false,
"language": [],
"downloads": 1513631,
"publisher": "banned-historical-archives",
"created_on": "2023-12-17T14:47:08.000Z",
"dataset_id": "huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
"dataset_url": "https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
"description": "和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.",
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [
"n<1K"
],
"task_categories": [],
"last_modified_on": "2025-10-19T15:21:40.000Z",
"observation_week": "2026-W27",
"card_task_categories": []
}05ksolovev/FineNews · 1,864,227 downloads1,864,227 downloads{"sha":"41f8a91dc8c493d0773bc0ebc5158886f5964924","tags":["region:us"],"gated":f…
EVENT. cmr2soi5ID. cmr2soi530515kh0czu19insvSRC. key:cmpxakb6…
{
"sha": "41f8a91dc8c493d0773bc0ebc5158886f5964924",
"tags": [
"region:us"
],
"gated": false,
"likes": 16,
"license": null,
"private": false,
"language": [],
"downloads": 1864227,
"publisher": "ksolovev",
"created_on": "2026-03-21T09:09:42.000Z",
"dataset_id": "huggingface.co/datasets/ksolovev/FineNews",
"dataset_url": "https://huggingface.co/datasets/ksolovev/FineNews",
"description": null,
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": null,
"card_license": null,
"card_languages": [],
"size_categories": [],
"task_categories": [],
"last_modified_on": "2026-03-23T23:30:10.000Z",
"observation_week": "2026-W27",
"card_task_categories": []
}06KakologArchives/KakologArchives · 1,898,322 downloadsニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA な{"sha":"27c129d9dc3be06c85c035c291f16d277b948f08","tags":["task_categories:text-…
EVENT. cmr2sohjID. cmr2sohjv0513kh0ccsd1l12sSRC. key:cmpxakb6…
{
"sha": "27c129d9dc3be06c85c035c291f16d277b948f08",
"tags": [
"task_categories:text-classification",
"language:ja",
"license:mit",
"region:us"
],
"gated": false,
"likes": 58,
"license": "mit",
"private": false,
"language": [
"ja"
],
"downloads": 1898322,
"publisher": "KakologArchives",
"created_on": "2023-05-12T13:31:56.000Z",
"dataset_id": "huggingface.co/datasets/KakologArchives/KakologArchives",
"dataset_url": "https://huggingface.co/datasets/KakologArchives/KakologArchives",
"description": "ニコニコ実況 過去ログアーカイブ ニコニコ実況 過去ログアーカイブは、ニコニコ実況 のサービス開始から現在までのすべての過去ログコメントを収集したデータセットです。 去る2020年12月、ニコニコ実況は ニコニコ生放送内の一公式チャンネルとしてリニューアル されました。これに伴い、2009年11月から運用されてきた旧システムは提供終了となり(事実上のサービス終了)、torne や BRAVIA などの家電への対応が軒並み終了する中、当時の生の声が詰まった約11年分の過去ログも同時に失われることとなってしまいました。 そこで 5ch の DTV 板の住民が中心となり、旧ニコニコ実況が終了するまでに11年分の全チャンネルの過去ログをアーカイブする計画が立ち上がりました。紆余曲折あり Nekopanda 氏が約11年分のラジオや BS も含めた全チャンネルの過去ログを完璧に取得してくださったおかげで、11年分の過去ログが電子の海に消えていく事態は回避できました。しかし、旧 API が廃止されてしまったため過去ログを API… See the full description on the dataset page: https://huggingface.co/datasets/KakologArchives/KakologArchives.",
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": "ニコニコ実況 過去ログアーカイブ",
"card_license": "mit",
"card_languages": [
"ja"
],
"size_categories": [],
"task_categories": [
"text-classification"
],
"last_modified_on": "2026-07-02T00:52:10.000Z",
"observation_week": "2026-W27",
"card_task_categories": [
"text-classification"
]
}07anisoleai/fineweb-tokenized · 2,130,303 downloadsFineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokeniz{"sha":"ef1311f460b42138d7a2d18f51e9cc38cedda089","tags":["task_categories:text-…
EVENT. cmr2sogyID. cmr2sogyu0511kh0cmfmyazvfSRC. key:cmpxakb6…
{
"sha": "ef1311f460b42138d7a2d18f51e9cc38cedda089",
"tags": [
"task_categories:text-generation",
"language:en",
"license:odc-by",
"size_categories:n>1T",
"modality:tabular",
"modality:text",
"arxiv:2406.17557",
"region:us",
"tabular",
"text",
"pre-training"
],
"gated": false,
"likes": 2,
"license": "odc-by",
"private": false,
"language": [
"en"
],
"downloads": 2130303,
"publisher": "anisoleai",
"created_on": "2026-05-26T06:34:59.000Z",
"dataset_id": "huggingface.co/datasets/anisoleai/fineweb-tokenized",
"dataset_url": "https://huggingface.co/datasets/anisoleai/fineweb-tokenized",
"description": "FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.",
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": "FineWeb Tokenized (AnisoleAI)",
"card_license": "odc-by",
"card_languages": [
"en"
],
"size_categories": [
"n>1T"
],
"task_categories": [
"text-generation"
],
"last_modified_on": "2026-05-29T12:35:53.000Z",
"observation_week": "2026-W27",
"card_task_categories": [
"text-generation"
]
}08Benjy/typed_digital_signatures · 2,141,013 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cmr2sogdID. cmr2sogdh050zkh0cmnnpq21zSRC. key:cmpxakb6…
{
"sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
"tags": [
"task_categories:image-classification",
"task_categories:zero-shot-image-classification",
"task_categories:image-feature-extraction",
"language:en",
"license:mit",
"size_categories:10K<n<100K",
"modality:image",
"region:us",
"digital-signatures",
"synthetic-data",
"image-classification",
"computer-vision",
"google-fonts",
"handwriting"
],
"gated": false,
"likes": 7,
"license": "mit",
"private": false,
"language": [
"en"
],
"downloads": 2141013,
"publisher": "Benjy",
"created_on": "2025-01-13T19:08:10.000Z",
"dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
"dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
"description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size: ~90,000… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": "Typed Digital Signatures Dataset",
"card_license": "mit",
"card_languages": [
"en"
],
"size_categories": [
"10K<n<100K"
],
"task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
],
"last_modified_on": "2025-02-05T18:44:13.000Z",
"observation_week": "2026-W27",
"card_task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
]
}09huggingface/documentation-images · 3,018,661 downloadsThis dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng{"sha":"c244738f5ae25d28b33bcee5ab25ccd3b66ae48d","tags":["license:cc-by-nc-sa-4…
EVENT. cmr2sofsID. cmr2sofs8050xkh0cuhkzd81vSRC. key:cmpxakb6…
{
"sha": "c244738f5ae25d28b33bcee5ab25ccd3b66ae48d",
"tags": [
"license:cc-by-nc-sa-4.0",
"size_categories:n<1K",
"format:imagefolder",
"modality:image",
"library:datasets",
"library:mlcroissant",
"region:us"
],
"gated": false,
"likes": 160,
"license": "cc-by-nc-sa-4.0",
"private": false,
"language": [],
"downloads": 3018661,
"publisher": "huggingface",
"created_on": "2022-03-02T23:29:22.000Z",
"dataset_id": "huggingface.co/datasets/huggingface/documentation-images",
"dataset_url": "https://huggingface.co/datasets/huggingface/documentation-images",
"description": "This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize the assets before uploading them. My favorite tool for this is https://tinypng.com/.",
"observed_at": "2026-07-02T00:57:14.439Z",
"pretty_name": null,
"card_license": "cc-by-nc-sa-4.0",
"card_languages": [],
"size_categories": [
"n<1K"
],
"task_categories": [],
"last_modified_on": "2026-07-01T13:27:12.000Z",
"observation_week": "2026-W27",
"card_task_categories": []
}10Benjy/typed_digital_signatures · 634,532 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cmqjj9v9ID. cmqjj9v9o0129mm0c7djwqbkoSRC. key:cmpxakb6…
{
"sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
"tags": [
"task_categories:image-classification",
"task_categories:zero-shot-image-classification",
"task_categories:image-feature-extraction",
"language:en",
"license:mit",
"size_categories:10K<n<100K",
"modality:image",
"region:us",
"digital-signatures",
"synthetic-data",
"image-classification",
"computer-vision",
"google-fonts",
"handwriting"
],
"gated": false,
"likes": 6,
"license": "mit",
"private": false,
"language": [
"en"
],
"downloads": 634532,
"publisher": "Benjy",
"created_on": "2025-01-13T19:08:10.000Z",
"dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
"dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
"description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size: ~90,000… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
"observed_at": "2026-06-18T13:26:20.060Z",
"pretty_name": "Typed Digital Signatures Dataset",
"card_license": "mit",
"card_languages": [
"en"
],
"size_categories": [
"10K<n<100K"
],
"task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
],
"last_modified_on": "2025-02-05T18:44:13.000Z",
"observation_week": "2026-W25",
"card_task_categories": [
"image-classification",
"zero-shot-image-classification",
"image-feature-extraction"
]
}showing 1–10 of 80older →
§03
subscribe
three pathways carry every event on this topic. pick the one that fits your agent.
GETrss feed
any reader · no authhttps://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xmlGETjson pull
poll on your schedule · optional since/untilhttps://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.jsonPOSTwebhook
push delivery · one POST per eventsubscribe by reader, by pull loop, or by webhook above