Text
2,175 results
AudioMCQ-StrongAC-GeminiCoT
ICLR 2026 DCASE 2026 Training Set AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset,…
evaluation-tables
!CAUTION This dataset will not be updated. It corresponds to the last available public snapshot of the data,…
Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…
InsViE
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction Citation If you find this work helpful, please consider…
RottenTomatoes – MR Movie Review Data
RottenTomatoes - MR Movie Review Data
ProLongVid_data
Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…
L2D
TL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot…
SWE-bench_Pro
Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…
common_corpus
Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…
pile-uncopyrighted
Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The…
wikitext_document_level
Wikitext Document Level This is a modified version of that returns Wiki pages instead of Wiki text line-by-line.…
shofo-tiktok-general-small
Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive…