New Datasets

topic · knowledge/datasets-published
DOC.
knowledge/datasets-published
REV.
179 evt
DATE.
02-JUN-2026
SCOPE.
custom
§01

about

Top-downloaded datasets on the Hugging Face Hub.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 152 events in this window (179 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01k9cli/video-vec2wav2-tokenizer · 909,987 downloadsvideo-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recogni{"sha":"97ef11e9bb9ef199c7e015551408609b90db6637","tags":["region:us"],"gated":f…
EVENT. cmrsii7rID. cmrsii7ry6zsrkh0cljyn4u0tSRC. key:cmpxakb6
{
  "sha": "97ef11e9bb9ef199c7e015551408609b90db6637",
  "tags": [
    "region:us"
  ],
  "gated": false,
  "likes": 10,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 909987,
  "publisher": "k9cli",
  "created_on": "2026-06-19T10:28:38.000Z",
  "dataset_id": "huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer",
  "dataset_url": "https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer",
  "description": "video-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-07-18T10:34:57.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
02openai/gsm8k · 925,692 downloadsDataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the tas{"sha":"740312add88f781978c0658806c59bc2815b9866","tags":["benchmark:official","…
EVENT. cmrsii79ID. cmrsii79p6zspkh0cz5kx5m0nSRC. key:cmpxakb6
{
  "sha": "740312add88f781978c0658806c59bc2815b9866",
  "tags": [
    "benchmark:official",
    "benchmark:eval-yaml",
    "task_categories:text-generation",
    "annotations_creators:crowdsourced",
    "language_creators:crowdsourced",
    "multilinguality:monolingual",
    "source_datasets:original",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:pandas",
    "library:polars",
    "library:mlcroissant",
    "arxiv:2110.14168",
    "region:us",
    "math-word-problems"
  ],
  "gated": false,
  "likes": 1449,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 925692,
  "publisher": "openai",
  "created_on": "2022-04-12T10:22:10.000Z",
  "dataset_id": "huggingface.co/datasets/openai/gsm8k",
  "dataset_url": "https://huggingface.co/datasets/openai/gsm8k",
  "description": "Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": "Grade School Math 8K",
  "card_license": [
    "mit"
  ],
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2026-03-23T10:18:13.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "text-generation"
  ]
}
03xlangai/ubuntu_osworld_file_cache · 1,111,684 downloadsOSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld {"sha":"711e0811642364e7aa8f10a8918367d0b626d578","tags":["license:apache-2.0","…
EVENT. cmrsii6rID. cmrsii6rc6zsnkh0cbvo4xdkkSRC. key:cmpxakb6
{
  "sha": "711e0811642364e7aa8f10a8918367d0b626d578",
  "tags": [
    "license:apache-2.0",
    "arxiv:2404.07972",
    "region:us"
  ],
  "gated": false,
  "likes": 40,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 1111684,
  "publisher": "xlangai",
  "created_on": "2025-05-27T15:19:04.000Z",
  "dataset_id": "huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache",
  "dataset_url": "https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache",
  "description": "OSWorld File Cache This repository serves as a file cache for the OSWorld project, providing reliable and fast access to evaluation files that were previously hosted on Google Drive. Overview OSWorld is a scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across various operating systems and applications. This cache repository ensures that all evaluation files are consistently… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/ubuntu_osworld_file_cache.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-05-07T11:32:14.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
04banned-historical-archives/banned-historical-archives · 1,247,317 downloads和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw{"sha":"7ee825405a889ac77b0d404fac868f9813da0c83","tags":["size_categories:n<1K"…
EVENT. cmrsii69ID. cmrsii6936zslkh0ctc65ioxcSRC. key:cmpxakb6
{
  "sha": "7ee825405a889ac77b0d404fac868f9813da0c83",
  "tags": [
    "size_categories:n<1K",
    "format:imagefolder",
    "modality:image",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 60,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 1247317,
  "publisher": "banned-historical-archives",
  "created_on": "2023-12-17T14:47:08.000Z",
  "dataset_id": "huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
  "dataset_url": "https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives",
  "description": "和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步 raw # 原始文件 config # 配置文件 todo # 存放暂未录入网站的文件 部分报纸和图片资料存放在单独的仓库: 名称 地址 状态 参考消息 https://huggingface.co/datasets/banned-historical-archives/ckxx 未录入 人民日报 https://huggingface.co/datasets/banned-historical-archives/rmrb 已精选重要的文章录入 文汇报… See the full description on the dataset page: https://huggingface.co/datasets/banned-historical-archives/banned-historical-archives.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2025-10-19T15:21:40.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
05Salesforce/wikitext · 1,326,422 downloadsDataset Card for "wikitext" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia{"sha":"b08601e04326c79dfdd32d625aee71d232d685c3","tags":["task_categories:text-…
EVENT. cmrsii5qID. cmrsii5qt6zsjkh0crth3w9n5SRC. key:cmpxakb6
{
  "sha": "b08601e04326c79dfdd32d625aee71d232d685c3",
  "tags": [
    "task_categories:text-generation",
    "task_categories:fill-mask",
    "task_ids:language-modeling",
    "task_ids:masked-language-modeling",
    "annotations_creators:no-annotation",
    "language_creators:crowdsourced",
    "multilinguality:monolingual",
    "source_datasets:original",
    "language:en",
    "license:cc-by-sa-3.0",
    "license:gfdl",
    "size_categories:1M<n<10M",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:dask",
    "library:polars",
    "library:mlcroissant",
    "arxiv:1609.07843",
    "region:us"
  ],
  "gated": false,
  "likes": 748,
  "license": "cc-by-sa-3.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 1326422,
  "publisher": "Salesforce",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/Salesforce/wikitext",
  "dataset_url": "https://huggingface.co/datasets/Salesforce/wikitext",
  "description": "Dataset Card for \"wikitext\" Dataset Summary The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia. The dataset is available under the Creative Commons Attribution-ShareAlike License. Compared to the preprocessed version of Penn Treebank (PTB), WikiText-2 is over 2 times larger and WikiText-103 is over 110 times larger. The WikiText dataset also features a far… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/wikitext.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": "WikiText",
  "card_license": [
    "cc-by-sa-3.0",
    "gfdl"
  ],
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "1M<n<10M"
  ],
  "task_categories": [
    "text-generation",
    "fill-mask"
  ],
  "last_modified_on": "2024-01-04T16:49:18.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "text-generation",
    "fill-mask"
  ]
}
06m-a-p/FineFineWeb · 1,383,211 downloadsFineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iterati{"sha":"7fd92dc825a75cbff271a5a52eea0eda91a2c112","tags":["task_categories:text-…
EVENT. cmrsii58ID. cmrsii58a6zshkh0cbikfkr4oSRC. key:cmpxakb6
{
  "sha": "7fd92dc825a75cbff271a5a52eea0eda91a2c112",
  "tags": [
    "task_categories:text-classification",
    "task_categories:text-generation",
    "language:en",
    "license:apache-2.0",
    "size_categories:1B<n<10B",
    "modality:tabular",
    "modality:text",
    "region:us"
  ],
  "gated": false,
  "likes": 160,
  "license": "apache-2.0",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 1383211,
  "publisher": "m-a-p",
  "created_on": "2024-12-14T12:46:33.000Z",
  "dataset_id": "huggingface.co/datasets/m-a-p/FineFineWeb",
  "dataset_url": "https://huggingface.co/datasets/m-a-p/FineFineWeb",
  "description": "FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "1B<n<10B"
  ],
  "task_categories": [
    "text-classification",
    "text-generation"
  ],
  "last_modified_on": "2024-12-19T11:34:03.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "text-classification",
    "text2text-generation",
    "text-generation"
  ]
}
07ayuo/hd_tmp · 1,399,124 downloads1,399,124 downloads{"sha":"0b7fd955a202c05b39cf9bc816c270fbdb391ea4","tags":["region:us"],"gated":f…
EVENT. cmrsii4qID. cmrsii4q06zsfkh0cx4ou6w9eSRC. key:cmpxakb6
{
  "sha": "0b7fd955a202c05b39cf9bc816c270fbdb391ea4",
  "tags": [
    "region:us"
  ],
  "gated": false,
  "likes": 25,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 1399124,
  "publisher": "ayuo",
  "created_on": "2026-04-01T14:22:50.000Z",
  "dataset_id": "huggingface.co/datasets/ayuo/hd_tmp",
  "dataset_url": "https://huggingface.co/datasets/ayuo/hd_tmp",
  "description": null,
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-07-05T09:46:14.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
08ryanmarten/OpenThoughts-1k-sample · 1,403,363 downloads[!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality {"sha":"a82400884621626d41bef89b7604f8054e7e00e0","tags":["size_categories:1K<n<…
EVENT. cmrsii47ID. cmrsii47p6zsdkh0ckhcbi7hnSRC. key:cmpxakb6
{
  "sha": "a82400884621626d41bef89b7604f8054e7e00e0",
  "tags": [
    "size_categories:1K<n<10K",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:pandas",
    "library:mlcroissant",
    "library:polars",
    "arxiv:2506.04178",
    "region:us"
  ],
  "gated": false,
  "likes": 41,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 1403363,
  "publisher": "ryanmarten",
  "created_on": "2025-08-30T23:58:46.000Z",
  "dataset_id": "huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample",
  "dataset_url": "https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample",
  "description": "[!NOTE] We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample of the OpenThoughts-114k dataset. Open synthetic reasoning dataset with high-quality examples covering math, science, code, and puzzles! Inspect the content with rich formatting with Curator Viewer. Available Subsets default subset containing ready-to-train data used to finetune the OpenThinker-7B and OpenThinker-32B models: ds =… See the full description on the dataset page: https://huggingface.co/datasets/ryanmarten/OpenThoughts-1k-sample.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "1K<n<10K"
  ],
  "task_categories": [],
  "last_modified_on": "2025-08-31T00:33:15.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
09hallucinations-leaderboard/results · 1,503,804 downloads1,503,804 downloads{"sha":"30f759225f89edcc18cb403b93f27bf4231223f7","tags":["license:apache-2.0","…
EVENT. cmrsii3pID. cmrsii3pc6zsbkh0cxg2aezo0SRC. key:cmpxakb6
{
  "sha": "30f759225f89edcc18cb403b93f27bf4231223f7",
  "tags": [
    "license:apache-2.0",
    "region:us"
  ],
  "gated": false,
  "likes": 14,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 1503804,
  "publisher": "hallucinations-leaderboard",
  "created_on": "2023-11-21T11:44:46.000Z",
  "dataset_id": "huggingface.co/datasets/hallucinations-leaderboard/results",
  "dataset_url": "https://huggingface.co/datasets/hallucinations-leaderboard/results",
  "description": null,
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2024-10-31T20:32:52.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": []
}
10allenai/c4 · 1,540,762 downloadsC4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We pre{"sha":"1588ec454efa1a09f29cd18ddd04fe05fc8653a2","tags":["task_categories:text-…
EVENT. cmrsii36ID. cmrsii36s6zs9kh0cimx4qacmSRC. key:cmpxakb6
{
  "sha": "1588ec454efa1a09f29cd18ddd04fe05fc8653a2",
  "tags": [
    "task_categories:text-generation",
    "task_categories:fill-mask",
    "task_ids:language-modeling",
    "task_ids:masked-language-modeling",
    "annotations_creators:no-annotation",
    "language_creators:found",
    "multilinguality:multilingual",
    "source_datasets:original",
    "language:af",
    "language:am",
    "language:ar",
    "language:az",
    "language:be",
    "language:bg",
    "language:bn",
    "language:ca",
    "language:ceb",
    "language:co",
    "language:cs",
    "language:cy",
    "language:da",
    "language:de",
    "language:el",
    "language:en",
    "language:eo",
    "language:es",
    "language:et",
    "language:eu",
    "language:fa",
    "language:fi"
  ],
  "gated": false,
  "likes": 618,
  "license": "odc-by",
  "private": false,
  "language": [
    "af",
    "am",
    "ar",
    "az",
    "be",
    "bg",
    "bn",
    "ca",
    "ceb",
    "co",
    "cs",
    "cy",
    "da",
    "de",
    "el",
    "en",
    "eo",
    "es",
    "et",
    "eu",
    "fa",
    "fi",
    "fil",
    "fr",
    "fy",
    "ga",
    "gd",
    "gl",
    "gu",
    "ha",
    "haw",
    "he",
    "hi",
    "hmn",
    "ht",
    "hu",
    "hy",
    "id",
    "ig",
    "is",
    "it",
    "iw",
    "ja",
    "jv",
    "ka",
    "kk",
    "km",
    "kn",
    "ko",
    "ku",
    "ky",
    "la",
    "lb",
    "lo",
    "lt",
    "lv",
    "mg",
    "mi",
    "mk",
    "ml",
    "mn",
    "mr",
    "ms",
    "mt",
    "my",
    "ne",
    "nl",
    "no",
    "ny",
    "pa",
    "pl",
    "ps",
    "pt",
    "ro",
    "ru",
    "sd",
    "si",
    "sk",
    "sl",
    "sm",
    "sn",
    "so",
    "sq",
    "sr",
    "st",
    "su",
    "sv",
    "sw",
    "ta",
    "te",
    "tg",
    "th",
    "tr",
    "uk",
    "und",
    "ur",
    "uz",
    "vi",
    "xh",
    "yi",
    "yo",
    "zh",
    "zu"
  ],
  "downloads": 1540762,
  "publisher": "allenai",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/allenai/c4",
  "dataset_url": "https://huggingface.co/datasets/allenai/c4",
  "description": "C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: \"https://commoncrawl.org\". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.",
  "observed_at": "2026-07-20T00:54:17.948Z",
  "pretty_name": "C4",
  "card_license": [
    "odc-by"
  ],
  "card_languages": [
    "af",
    "am",
    "ar",
    "az",
    "be"
  ],
  "size_categories": [
    "10B<n<100B"
  ],
  "task_categories": [
    "text-generation",
    "fill-mask"
  ],
  "last_modified_on": "2024-01-09T19:14:03.000Z",
  "observation_week": "2026-W30",
  "card_task_categories": [
    "text-generation",
    "fill-mask"
  ]
}
showing 1–10 of 152older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above