Skip to content
Advertisement

Tabular

1,117 results

MultimodalTabularText

gaming-500-hours

Gaming Dataset (gaming-1) — 494.7 Hours Native PC/console gameplay screen-recordings, organized by game. Each workflow is one play…

<1K·JSON
MultimodalTabularText

smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…

100M–1B·ODC-BY·Parquet
MultimodalTabularText

Goodreads-Books

Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…

1M–10M·Custom / Research-only·CSV
MultimodalTabularText

common_corpus

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…

10K–100K·Parquet
Advertisement