New Datasets

topic · knowledge/datasets-published
DOC.
knowledge/datasets-published
REV.
179 evt
DATE.
02-JUN-2026
SCOPE.
custom
§01

about

Top-downloaded datasets on the Hugging Face Hub.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 116 events in this window (179 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01Benjy/typed_digital_signatures · 3,109,722 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cmriihd3ID. cmriihd3l4ctbkh0c2nzjkhuvSRC. key:cmpxakb6
{
  "sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
  "tags": [
    "task_categories:image-classification",
    "task_categories:zero-shot-image-classification",
    "task_categories:image-feature-extraction",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "modality:image",
    "region:us",
    "digital-signatures",
    "synthetic-data",
    "image-classification",
    "computer-vision",
    "google-fonts",
    "handwriting"
  ],
  "gated": false,
  "likes": 11,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 3109722,
  "publisher": "Benjy",
  "created_on": "2025-01-13T19:08:10.000Z",
  "dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
  "dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
  "description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size:… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": "Typed Digital Signatures Dataset",
  "card_license": "mit",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "image-classification",
    "zero-shot-image-classification",
    "image-feature-extraction"
  ],
  "last_modified_on": "2025-02-05T18:44:13.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "image-classification",
    "zero-shot-image-classification",
    "image-feature-extraction"
  ]
}
02codeparrot/github-code · 5,702,880 downloadsThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuer{"sha":"b5661e6b17396364b2bcf8e68977b0d28e1ebd19","tags":["task_categories:text-…
EVENT. cmriihckID. cmriihcks4ct9kh0cf5ql8mgbSRC. key:cmpxakb6
{
  "sha": "b5661e6b17396364b2bcf8e68977b0d28e1ebd19",
  "tags": [
    "task_categories:text-generation",
    "task_ids:language-modeling",
    "language_creators:crowdsourced",
    "language_creators:expert-generated",
    "multilinguality:multilingual",
    "language:code",
    "license:other",
    "region:us"
  ],
  "gated": false,
  "likes": 379,
  "license": "other",
  "private": false,
  "language": [
    "code"
  ],
  "downloads": 5702880,
  "publisher": "codeparrot",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/codeparrot/github-code",
  "dataset_url": "https://huggingface.co/datasets/codeparrot/github-code",
  "description": "The GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": "github-code",
  "card_license": [
    "other"
  ],
  "card_languages": [
    "code"
  ],
  "size_categories": [],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2022-10-20T15:01:14.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "text-generation"
  ]
}
03anisoleai/fineweb-tokenized · 6,075,318 downloadsFineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokeniz{"sha":"ef1311f460b42138d7a2d18f51e9cc38cedda089","tags":["task_categories:text-…
EVENT. cmriihbuID. cmriihbu74ct5kh0chpgnm0aoSRC. key:cmpxakb6
{
  "sha": "ef1311f460b42138d7a2d18f51e9cc38cedda089",
  "tags": [
    "task_categories:text-generation",
    "language:en",
    "license:odc-by",
    "size_categories:n>1T",
    "modality:tabular",
    "modality:text",
    "arxiv:2406.17557",
    "region:us",
    "tabular",
    "text",
    "pre-training"
  ],
  "gated": false,
  "likes": 8,
  "license": "odc-by",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 6075318,
  "publisher": "anisoleai",
  "created_on": "2026-05-26T06:34:59.000Z",
  "dataset_id": "huggingface.co/datasets/anisoleai/fineweb-tokenized",
  "dataset_url": "https://huggingface.co/datasets/anisoleai/fineweb-tokenized",
  "description": "FineWeb Tokenized > 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page: https://huggingface.co/datasets/anisoleai/fineweb-tokenized.",
  "observed_at": "2026-07-13T00:56:05.317Z",
  "pretty_name": "FineWeb Tokenized (AnisoleAI)",
  "card_license": "odc-by",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2026-05-29T12:35:53.000Z",
  "observation_week": "2026-W29",
  "card_task_categories": [
    "text-generation"
  ]
}
04mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M · 735,591 downloads🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀 Upload Status All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics 📜 Cite If you find LLaVA-One-V{"sha":"c5218cad785eba7d218137e8ce4997bda568a050","tags":["license:apache-2.0","…
EVENT. cmrdvdwwID. cmrdvdwwu356hkh0cq3neslp5SRC. key:cmpxakb6
{
  "sha": "c5218cad785eba7d218137e8ce4997bda568a050",
  "tags": [
    "license:apache-2.0",
    "size_categories:10M<n<100M",
    "format:parquet",
    "modality:image",
    "modality:text",
    "library:datasets",
    "library:dask",
    "library:mlcroissant",
    "library:polars",
    "arxiv:2509.23661",
    "region:us"
  ],
  "gated": false,
  "likes": 75,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 735591,
  "publisher": "mvp-lab",
  "created_on": "2025-09-14T14:42:33.000Z",
  "dataset_id": "huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M",
  "dataset_url": "https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M",
  "description": "🚀 LLaVA-One-Vision-1.5-Mid-Training-85M Dataset is being uploaded 🚀 Upload Status All Completed: ImageNet-21k、LAIONCN、DataComp-1B、Zero250M、COYO700M、SA-1B、MINT、Obelics 📜 Cite If you find LLaVA-One-Vision-1.5-Mid-Training-85M useful in your research, please consider to cite the following related papers: @misc{an2025llavaonevision15fullyopenframework, title={LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training}… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Mid-Training-85M.",
  "observed_at": "2026-07-09T18:58:30.193Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [
    "10M<n<100M"
  ],
  "task_categories": [],
  "last_modified_on": "2025-11-24T06:32:02.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": []
}
05nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim · 745,992 downloadsPhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robo{"sha":"ea7ac0b68f87da62f1e726771bba0fe74300802f","tags":["task_categories:robot…
EVENT. cmr984q9ID. cmr984q9k1vj1kh0c9sze49y4SRC. key:cmpxakb6
{
  "sha": "ea7ac0b68f87da62f1e726771bba0fe74300802f",
  "tags": [
    "task_categories:robotics",
    "license:cc-by-4.0",
    "region:us",
    "robotics"
  ],
  "gated": false,
  "likes": 239,
  "license": "cc-by-4.0",
  "private": false,
  "language": [],
  "downloads": 745992,
  "publisher": "nvidia",
  "created_on": "2025-03-18T13:59:39.000Z",
  "dataset_id": "huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "dataset_url": "https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim",
  "description": "PhysicalAI-Robotics-GR00T-X-Embodiment-Sim Github Repo: Isaac GR00T N1 We provide a set of datasets used for post-training of GR00T N1. Each dataset is a collection of trajectories from different robot embodiments and tasks. Cross-embodied bimanual manipulation: 9k trajectories Dataset Name #trajectories bimanual_panda_gripper.Threading 1000 bimanual_panda_hand.LiftTray 1000 bimanual_panda_gripper.ThreePieceAssembly 1000… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-X-Embodiment-Sim.",
  "observed_at": "2026-07-06T12:56:25.781Z",
  "pretty_name": null,
  "card_license": "cc-by-4.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2026-03-05T23:36:40.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": [
    "robotics"
  ]
}
06HennyPr/ps2_hf2 · 762,549 downloadsn<1K · 762,549 downloads{"sha":"e831be1a0eeb18dbfd99aac845da6bb4c271d08b","tags":["size_categories:n<1K"…
EVENT. cmr8ie7yID. cmr8ie7y01orvkh0cvcjmzvjhSRC. key:cmpxakb6
{
  "sha": "e831be1a0eeb18dbfd99aac845da6bb4c271d08b",
  "tags": [
    "size_categories:n<1K",
    "format:text",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 21,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 762549,
  "publisher": "HennyPr",
  "created_on": "2026-03-10T03:07:26.000Z",
  "dataset_id": "huggingface.co/datasets/HennyPr/ps2_hf2",
  "dataset_url": "https://huggingface.co/datasets/HennyPr/ps2_hf2",
  "description": null,
  "observed_at": "2026-07-06T00:55:46.488Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-04-05T20:00:04.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": []
}
07osv5m/osv5m · 763,695 downloadsOpenStreetView-5M The Many Roads to Global Visual Geolocation 📍🌍 First authors: Guillaume Astruc, Nicolas Dufour, Ioannis SiglidisSecond authors: Constantin Aronssohn, Nacim Bouia, Stephanie Fu, Rom{"sha":"cff33609b56b54d8743b7ee7a416eb8433e9a681","tags":["license:cc-by-sa-4.0"…
EVENT. cmr8ie7gID. cmr8ie7gl1ortkh0cunuw9ic4SRC. key:cmpxakb6
{
  "sha": "cff33609b56b54d8743b7ee7a416eb8433e9a681",
  "tags": [
    "license:cc-by-sa-4.0",
    "region:us"
  ],
  "gated": false,
  "likes": 56,
  "license": "cc-by-sa-4.0",
  "private": false,
  "language": [],
  "downloads": 763695,
  "publisher": "osv5m",
  "created_on": "2023-12-20T13:16:08.000Z",
  "dataset_id": "huggingface.co/datasets/osv5m/osv5m",
  "dataset_url": "https://huggingface.co/datasets/osv5m/osv5m",
  "description": "OpenStreetView-5M The Many Roads to Global Visual Geolocation 📍🌍 First authors: Guillaume Astruc, Nicolas Dufour, Ioannis SiglidisSecond authors: Constantin Aronssohn, Nacim Bouia, Stephanie Fu, Romain Loiseau, Van Nguyen Nguyen, Charles Raude, Elliot Vincent, Lintao XU, Hongyu ZhouLast author: Loic LandrieuResearch Institute: Imagine, LIGM, Ecole des Ponts, Univ Gustave Eiffel, CNRS, Marne-la-Vallée, France Introduction 🌍 OpenStreetView-5M is the first large-scale… See the full description on the dataset page: https://huggingface.co/datasets/osv5m/osv5m.",
  "observed_at": "2026-07-06T00:55:46.488Z",
  "pretty_name": null,
  "card_license": "cc-by-sa-4.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2024-04-27T01:39:53.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": []
}
08daniilakk/nbchr_pdfs · 918,267 downloads918,267 downloads{"sha":"2f410c2bae0caccd5297d3492d663bea4bbf2bda","tags":["license:unknown","mod…
EVENT. cmr8ie6yID. cmr8ie6yz1orrkh0cw70p3ur2SRC. key:cmpxakb6
{
  "sha": "2f410c2bae0caccd5297d3492d663bea4bbf2bda",
  "tags": [
    "license:unknown",
    "modality:document",
    "region:us"
  ],
  "gated": false,
  "likes": 19,
  "license": "unknown",
  "private": false,
  "language": [],
  "downloads": 918267,
  "publisher": "daniilakk",
  "created_on": "2025-05-24T09:47:35.000Z",
  "dataset_id": "huggingface.co/datasets/daniilakk/nbchr_pdfs",
  "dataset_url": "https://huggingface.co/datasets/daniilakk/nbchr_pdfs",
  "description": null,
  "observed_at": "2026-07-06T00:55:46.488Z",
  "pretty_name": null,
  "card_license": "unknown",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2025-05-25T05:00:50.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": []
}
09openai/gsm8k · 947,424 downloadsDataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the tas{"sha":"740312add88f781978c0658806c59bc2815b9866","tags":["benchmark:official","…
EVENT. cmr8ie6hID. cmr8ie6hl1orpkh0c58gtudjqSRC. key:cmpxakb6
{
  "sha": "740312add88f781978c0658806c59bc2815b9866",
  "tags": [
    "benchmark:official",
    "benchmark:eval-yaml",
    "task_categories:text-generation",
    "annotations_creators:crowdsourced",
    "language_creators:crowdsourced",
    "multilinguality:monolingual",
    "source_datasets:original",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:pandas",
    "library:polars",
    "library:mlcroissant",
    "arxiv:2110.14168",
    "region:us",
    "math-word-problems"
  ],
  "gated": false,
  "likes": 1419,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 947424,
  "publisher": "openai",
  "created_on": "2022-04-12T10:22:10.000Z",
  "dataset_id": "huggingface.co/datasets/openai/gsm8k",
  "dataset_url": "https://huggingface.co/datasets/openai/gsm8k",
  "description": "Dataset Card for GSM8K Dataset Summary GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning. These problems take between 2 and 8 steps to solve. Solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the… See the full description on the dataset page: https://huggingface.co/datasets/openai/gsm8k.",
  "observed_at": "2026-07-06T00:55:46.488Z",
  "pretty_name": "Grade School Math 8K",
  "card_license": [
    "mit"
  ],
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "text-generation"
  ],
  "last_modified_on": "2026-03-23T10:18:13.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": [
    "text-generation"
  ]
}
10genrobot2025/10Kh-RealOmin-OpenData · 1,205,822 downloadsBoasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,00{"sha":"fcbc0d38550e134f273426aa7c9cc2b491270bc4","tags":["task_categories:robot…
EVENT. cmr8ie60ID. cmr8ie60c1ornkh0ct83p03k6SRC. key:cmpxakb6
{
  "sha": "fcbc0d38550e134f273426aa7c9cc2b491270bc4",
  "tags": [
    "task_categories:robotics",
    "task_categories:reinforcement-learning",
    "language:en",
    "language:zh",
    "license:cc-by-sa-4.0",
    "size_categories:n>1T",
    "modality:video",
    "region:us",
    "agent",
    "robotic",
    "real-world",
    "dual-arm",
    "video",
    "vla",
    "embodied intelligence"
  ],
  "gated": "auto",
  "likes": 222,
  "license": "cc-by-sa-4.0",
  "private": false,
  "language": [
    "en",
    "zh"
  ],
  "downloads": 1205822,
  "publisher": "genrobot2025",
  "created_on": "2025-12-31T07:17:19.000Z",
  "dataset_id": "huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "dataset_url": "https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "description": "Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.",
  "observed_at": "2026-07-06T00:55:46.488Z",
  "pretty_name": null,
  "card_license": "cc-by-sa-4.0",
  "card_languages": [
    "en",
    "zh"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics",
    "reinforcement-learning"
  ],
  "last_modified_on": "2026-04-24T05:02:26.000Z",
  "observation_week": "2026-W28",
  "card_task_categories": [
    "robotics",
    "reinforcement-learning"
  ]
}
showing 1–10 of 116older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above