New Datasets

topic · knowledge/datasets-published
DOC.
knowledge/datasets-published
REV.
179 evt
DATE.
02-JUN-2026
SCOPE.
custom
§01

about

Top-downloaded datasets on the Hugging Face Hub.

§02

recent events

LIVElast event 0s ago0 evt / 1h

showing 10 of 71 events in this window (179 total on topic). adjust the range or clear it with ALL.

range
iso 8601 utc
iso 8601 utc
01Benjy/typed_digital_signatures · 634,532 downloadsTyped Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-s{"sha":"b2b7e3766c7b1a39430d44b9dbb96e915c819332","tags":["task_categories:image…
EVENT. cmqjj9v9ID. cmqjj9v9o0129mm0c7djwqbkoSRC. key:cmpxakb6
{
  "sha": "b2b7e3766c7b1a39430d44b9dbb96e915c819332",
  "tags": [
    "task_categories:image-classification",
    "task_categories:zero-shot-image-classification",
    "task_categories:image-feature-extraction",
    "language:en",
    "license:mit",
    "size_categories:10K<n<100K",
    "modality:image",
    "region:us",
    "digital-signatures",
    "synthetic-data",
    "image-classification",
    "computer-vision",
    "google-fonts",
    "handwriting"
  ],
  "gated": false,
  "likes": 6,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 634532,
  "publisher": "Benjy",
  "created_on": "2025-01-13T19:08:10.000Z",
  "dataset_id": "huggingface.co/datasets/Benjy/typed_digital_signatures",
  "dataset_url": "https://huggingface.co/datasets/Benjy/typed_digital_signatures",
  "description": "Typed Digital Signatures Dataset This comprehensive dataset contains synthetic digital signatures rendered across 30 different Google Fonts, specifically selected for their handwriting and signature-style characteristics. Each font contributes unique stylistic elements, making this dataset ideal for robust signature analysis and font recognition tasks. Dataset Overview Total Fonts: 30 different Google Fonts Images per Font: 3,000 signatures Total Dataset Size: ~90,000… See the full description on the dataset page: https://huggingface.co/datasets/Benjy/typed_digital_signatures.",
  "observed_at": "2026-06-18T13:26:20.060Z",
  "pretty_name": "Typed Digital Signatures Dataset",
  "card_license": "mit",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "10K<n<100K"
  ],
  "task_categories": [
    "image-classification",
    "zero-shot-image-classification",
    "image-feature-extraction"
  ],
  "last_modified_on": "2025-02-05T18:44:13.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": [
    "image-classification",
    "zero-shot-image-classification",
    "image-feature-extraction"
  ]
}
02genrobot2025/10Kh-RealOmin-OpenData · 876,332 downloadsBoasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,00{"sha":"fcbc0d38550e134f273426aa7c9cc2b491270bc4","tags":["task_categories:robot…
EVENT. cmqjj9upID. cmqjj9upe0127mm0cksvd5o04SRC. key:cmpxakb6
{
  "sha": "fcbc0d38550e134f273426aa7c9cc2b491270bc4",
  "tags": [
    "task_categories:robotics",
    "task_categories:reinforcement-learning",
    "language:en",
    "language:zh",
    "license:cc-by-sa-4.0",
    "size_categories:n>1T",
    "modality:video",
    "region:us",
    "agent",
    "robotic",
    "real-world",
    "dual-arm",
    "video",
    "vla",
    "embodied intelligence"
  ],
  "gated": "auto",
  "likes": 217,
  "license": "cc-by-sa-4.0",
  "private": false,
  "language": [
    "en",
    "zh"
  ],
  "downloads": 876332,
  "publisher": "genrobot2025",
  "created_on": "2025-12-31T07:17:19.000Z",
  "dataset_id": "huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "dataset_url": "https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData",
  "description": "Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.",
  "observed_at": "2026-06-18T13:26:20.060Z",
  "pretty_name": null,
  "card_license": "cc-by-sa-4.0",
  "card_languages": [
    "en",
    "zh"
  ],
  "size_categories": [
    "n>1T"
  ],
  "task_categories": [
    "robotics",
    "reinforcement-learning"
  ],
  "last_modified_on": "2026-04-24T05:02:26.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": [
    "robotics",
    "reinforcement-learning"
  ]
}
03SWE-bench/SWE-bench_Multilingual · 620,458 downloadsn<1K · 620,458 downloads{"sha":"2b7aced941b4873e9cad3e76abbae93f481d1beb","tags":["language:en","license…
EVENT. cmqehfp6ID. cmqehfp6i25rbnq0c8fqas6hoSRC. key:cmpxakb6
{
  "sha": "2b7aced941b4873e9cad3e76abbae93f481d1beb",
  "tags": [
    "language:en",
    "license:mit",
    "size_categories:n<1K",
    "format:parquet",
    "modality:text",
    "library:datasets",
    "library:pandas",
    "library:mlcroissant",
    "library:polars",
    "region:us"
  ],
  "gated": false,
  "likes": 17,
  "license": "mit",
  "private": false,
  "language": [
    "en"
  ],
  "downloads": 620458,
  "publisher": "SWE-bench",
  "created_on": "2025-04-29T00:37:54.000Z",
  "dataset_id": "huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual",
  "dataset_url": "https://huggingface.co/datasets/SWE-bench/SWE-bench_Multilingual",
  "description": null,
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": null,
  "card_license": "mit",
  "card_languages": [
    "en"
  ],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2025-08-26T00:42:28.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": []
}
04hf-doc-build/doc-build-dev · 649,297 downloadsThis is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider rep{"sha":"494fce0367580237ed1073a60ac801abb072f521","tags":["license:mit","region:…
EVENT. cmqehfokID. cmqehfokt25r7nq0ce8aim6eeSRC. key:cmpxakb6
{
  "sha": "494fce0367580237ed1073a60ac801abb072f521",
  "tags": [
    "license:mit",
    "region:us",
    "documentation"
  ],
  "gated": false,
  "likes": 36,
  "license": "mit",
  "private": false,
  "language": [],
  "downloads": 649297,
  "publisher": "hf-doc-build",
  "created_on": "2022-11-08T09:03:37.000Z",
  "dataset_id": "huggingface.co/datasets/hf-doc-build/doc-build-dev",
  "dataset_url": "https://huggingface.co/datasets/hf-doc-build/doc-build-dev",
  "description": "This is a dataset which contains the docs from all the PRs that are updating one of the docs from https://huggingface.co/docs. It is automatically updated by this github action from the doc-buider repo.",
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": "HF Documentation (PRs)",
  "card_license": "mit",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [],
  "last_modified_on": "2026-04-20T08:59:22.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": []
}
05jat-project/jat-dataset · 651,710 downloadsJAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textu{"sha":"921defdd6f8ae5271f2d4925150c48dba0cb3199","tags":["task_categories:reinf…
EVENT. cmqehfo1ID. cmqehfo1d25r3nq0c899eojwfSRC. key:cmpxakb6
{
  "sha": "921defdd6f8ae5271f2d4925150c48dba0cb3199",
  "tags": [
    "task_categories:reinforcement-learning",
    "task_categories:text-generation",
    "task_categories:question-answering",
    "annotations_creators:found",
    "annotations_creators:machine-generated",
    "source_datasets:conceptual-captions",
    "source_datasets:ok-vqa",
    "source_datasets:oscar",
    "license:apache-2.0",
    "size_categories:100M<n<1B",
    "format:parquet",
    "modality:image",
    "modality:text",
    "modality:timeseries",
    "library:datasets",
    "library:dask",
    "library:mlcroissant",
    "library:polars",
    "arxiv:2402.09844",
    "arxiv:2303.03915",
    "region:us",
    "imitation-learning",
    "reinforcement-learning",
    "text-generation",
    "question-answering",
    "generalist-agent"
  ],
  "gated": false,
  "likes": 60,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 651710,
  "publisher": "jat-project",
  "created_on": "2023-08-29T09:03:24.000Z",
  "dataset_id": "huggingface.co/datasets/jat-project/jat-dataset",
  "dataset_url": "https://huggingface.co/datasets/jat-project/jat-dataset",
  "description": "JAT Dataset Dataset Description The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent. Paper: https://huggingface.co/papers/2402.09844 Usage >>> from datasets import load_dataset >>> dataset = load_dataset(\"jat-project/jat-dataset\"… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.",
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": "JAT-dataset",
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [
    "100M<n<1B"
  ],
  "task_categories": [
    "reinforcement-learning",
    "text-generation",
    "question-answering"
  ],
  "last_modified_on": "2024-02-16T13:52:52.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": [
    "reinforcement-learning",
    "text-generation",
    "question-answering"
  ]
}
06Maynor996/img_upload · 660,537 downloadsn<1K · 660,537 downloads{"sha":"329267940d38d38a88a3a63a6e319075c7cfaf15","tags":["size_categories:n<1K"…
EVENT. cmqehfniID. cmqehfni925qznq0ci296r7gsSRC. key:cmpxakb6
{
  "sha": "329267940d38d38a88a3a63a6e319075c7cfaf15",
  "tags": [
    "size_categories:n<1K",
    "format:imagefolder",
    "modality:image",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 19,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 660537,
  "publisher": "Maynor996",
  "created_on": "2026-01-19T05:28:47.000Z",
  "dataset_id": "huggingface.co/datasets/Maynor996/img_upload",
  "dataset_url": "https://huggingface.co/datasets/Maynor996/img_upload",
  "description": null,
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-06-12T09:03:52.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": []
}
07Maynor996/upload2 · 722,879 downloadsn<1K · 722,879 downloads{"sha":"dadfbd3480c820e935f9e1715fd4fc9a2694894b","tags":["size_categories:n<1K"…
EVENT. cmqehfmzID. cmqehfmz525qvnq0c40iynckzSRC. key:cmpxakb6
{
  "sha": "dadfbd3480c820e935f9e1715fd4fc9a2694894b",
  "tags": [
    "size_categories:n<1K",
    "format:imagefolder",
    "modality:image",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 23,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 722879,
  "publisher": "Maynor996",
  "created_on": "2026-01-19T05:32:46.000Z",
  "dataset_id": "huggingface.co/datasets/Maynor996/upload2",
  "dataset_url": "https://huggingface.co/datasets/Maynor996/upload2",
  "description": null,
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-06-12T09:28:24.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": []
}
08IPEC-COMMUNITY/bridge_orig_lerobot · 734,003 downloadsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "widowx", "total_episodes": 53192, "total_frames": 1893026, "total_tasks": 19974, {"sha":"0e9d76d07e9df3ea3eba257b2520d4913833fad2","tags":["task_categories:robot…
EVENT. cmqehfmgID. cmqehfmgl25qrnq0cv8zzvkn4SRC. key:cmpxakb6
{
  "sha": "0e9d76d07e9df3ea3eba257b2520d4913833fad2",
  "tags": [
    "task_categories:robotics",
    "license:apache-2.0",
    "modality:video",
    "region:us",
    "LeRobot",
    "bridge_orig",
    "rlds",
    "openx",
    "widowx"
  ],
  "gated": false,
  "likes": 21,
  "license": "apache-2.0",
  "private": false,
  "language": [],
  "downloads": 734003,
  "publisher": "IPEC-COMMUNITY",
  "created_on": "2025-02-22T11:43:08.000Z",
  "dataset_id": "huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot",
  "dataset_url": "https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot",
  "description": "This dataset was created using LeRobot. Dataset Structure meta/info.json: { \"codebase_version\": \"v2.0\", \"robot_type\": \"widowx\", \"total_episodes\": 53192, \"total_frames\": 1893026, \"total_tasks\": 19974, \"total_videos\": 212768, \"total_chunks\": 54, \"chunks_size\": 1000, \"fps\": 5, \"splits\": { \"train\": \"0:53192\" }, \"data_path\": \"data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet\", \"video_path\":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bridge_orig_lerobot.",
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": null,
  "card_license": "apache-2.0",
  "card_languages": [],
  "size_categories": [],
  "task_categories": [
    "robotics"
  ],
  "last_modified_on": "2025-02-23T06:25:52.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": [
    "robotics"
  ]
}
09HennyPr/ps2_hf2 · 760,675 downloadsn<1K · 760,675 downloads{"sha":"e831be1a0eeb18dbfd99aac845da6bb4c271d08b","tags":["size_categories:n<1K"…
EVENT. cmqehflxID. cmqehflxu25qnnq0cvc9jjjxgSRC. key:cmpxakb6
{
  "sha": "e831be1a0eeb18dbfd99aac845da6bb4c271d08b",
  "tags": [
    "size_categories:n<1K",
    "format:text",
    "modality:text",
    "library:datasets",
    "library:mlcroissant",
    "region:us"
  ],
  "gated": false,
  "likes": 18,
  "license": null,
  "private": false,
  "language": [],
  "downloads": 760675,
  "publisher": "HennyPr",
  "created_on": "2026-03-10T03:07:26.000Z",
  "dataset_id": "huggingface.co/datasets/HennyPr/ps2_hf2",
  "dataset_url": "https://huggingface.co/datasets/HennyPr/ps2_hf2",
  "description": null,
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": null,
  "card_license": null,
  "card_languages": [],
  "size_categories": [
    "n<1K"
  ],
  "task_categories": [],
  "last_modified_on": "2026-04-05T20:00:04.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": []
}
10allenai/c4 · 826,781 downloadsC4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We pre{"sha":"1588ec454efa1a09f29cd18ddd04fe05fc8653a2","tags":["task_categories:text-…
EVENT. cmqehfleID. cmqehfleh25qjnq0ce1s4as3pSRC. key:cmpxakb6
{
  "sha": "1588ec454efa1a09f29cd18ddd04fe05fc8653a2",
  "tags": [
    "task_categories:text-generation",
    "task_categories:fill-mask",
    "task_ids:language-modeling",
    "task_ids:masked-language-modeling",
    "annotations_creators:no-annotation",
    "language_creators:found",
    "multilinguality:multilingual",
    "source_datasets:original",
    "language:af",
    "language:am",
    "language:ar",
    "language:az",
    "language:be",
    "language:bg",
    "language:bn",
    "language:ca",
    "language:ceb",
    "language:co",
    "language:cs",
    "language:cy",
    "language:da",
    "language:de",
    "language:el",
    "language:en",
    "language:eo",
    "language:es",
    "language:et",
    "language:eu",
    "language:fa",
    "language:fi"
  ],
  "gated": false,
  "likes": 597,
  "license": "odc-by",
  "private": false,
  "language": [
    "af",
    "am",
    "ar",
    "az",
    "be",
    "bg",
    "bn",
    "ca",
    "ceb",
    "co",
    "cs",
    "cy",
    "da",
    "de",
    "el",
    "en",
    "eo",
    "es",
    "et",
    "eu",
    "fa",
    "fi",
    "fil",
    "fr",
    "fy",
    "ga",
    "gd",
    "gl",
    "gu",
    "ha",
    "haw",
    "he",
    "hi",
    "hmn",
    "ht",
    "hu",
    "hy",
    "id",
    "ig",
    "is",
    "it",
    "iw",
    "ja",
    "jv",
    "ka",
    "kk",
    "km",
    "kn",
    "ko",
    "ku",
    "ky",
    "la",
    "lb",
    "lo",
    "lt",
    "lv",
    "mg",
    "mi",
    "mk",
    "ml",
    "mn",
    "mr",
    "ms",
    "mt",
    "my",
    "ne",
    "nl",
    "no",
    "ny",
    "pa",
    "pl",
    "ps",
    "pt",
    "ro",
    "ru",
    "sd",
    "si",
    "sk",
    "sl",
    "sm",
    "sn",
    "so",
    "sq",
    "sr",
    "st",
    "su",
    "sv",
    "sw",
    "ta",
    "te",
    "tg",
    "th",
    "tr",
    "uk",
    "und",
    "ur",
    "uz",
    "vi",
    "xh",
    "yi",
    "yo",
    "zh",
    "zu"
  ],
  "downloads": 826781,
  "publisher": "allenai",
  "created_on": "2022-03-02T23:29:22.000Z",
  "dataset_id": "huggingface.co/datasets/allenai/c4",
  "dataset_url": "https://huggingface.co/datasets/allenai/c4",
  "description": "C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: \"https://commoncrawl.org\". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one per… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.",
  "observed_at": "2026-06-15T00:35:48.674Z",
  "pretty_name": "C4",
  "card_license": [
    "odc-by"
  ],
  "card_languages": [
    "af",
    "am",
    "ar",
    "az",
    "be"
  ],
  "size_categories": [
    "10B<n<100B"
  ],
  "task_categories": [
    "text-generation",
    "fill-mask"
  ],
  "last_modified_on": "2024-01-09T19:14:03.000Z",
  "observation_week": "2026-W25",
  "card_task_categories": [
    "text-generation",
    "fill-mask"
  ]
}
showing 1–10 of 71older →
§03

subscribe

three pathways carry every event on this topic. pick the one that fits your agent.

GETrss feed
any reader · no auth
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published/feed.xml
GETjson pull
poll on your schedule · optional since/until
https://api.callsign.sh/v1/public/channels/knowledge/topics/datasets-published.json
POSTwebhook
push delivery · one POST per event
log in to subscribe →
subscribe by reader, by pull loop, or by webhook above