fineweb_deduplicated
TL;DR Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb…
1,501 results
TL;DR Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb…
The Join A broad collection of 650 relational databases spanning many domains (academic, e-commerce, finance, sports, biomedical, government,…
Carbon Pretraining Corpus
In-The-Wild Jailbreak Prompts on LLMs This is the official repository for the ACM CCS 2024 paper "Do Anything…
Game of 24 Mathematical Puzzle Dataset
Vietnamese Án lệ + Bản án Corpus
Changelog NEW Changes March 11th 2026 Added new split: arxiv papers, sourced from the Hugging Face /api/papers endpoint…
MMLU-ProX MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to…
Foursquare OS Places is now a gated dataset on Hugging Face. Read more about why we are making…
Betty Dota 2 — Decision Context Dataset Overview 9,385 professional Dota 2 matches parsed from replay files (.dem)…
NoeFlandre/osm-polygon-wikidata-only OSM polygons tagged with a wikidata= reference, enriched with Wikipedia and Wikivoyage text across all available…
CTU relational datasets (redelex) Relational databases from the CTU Prague Relational Learning Repository (a.k.a. the CTU relational repository),…
Dataset Card for "pickapic v2" please pay attention - the URLs will be temporariliy unavailabe - but you…
WhestBench 2026: ARC White-Box Estimation Challenge
Dataset Card for "QM9" QM9 dataset from Ruddigkeit et al., 2012; Ramakrishnan et al., 2014. Original data downloaded…
cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…
Crystallography Open Database (COD) — Full Snapshot A complete mirror of the Crystallography Open Database (COD) as a…
Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.…
Dataset Card for CodeForces-CoTs Dataset description CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive…