Skip to content
Advertisement

Text

2,175 results

Text

C-Eval

C-Eval

10K–100K·CC-BY-NC-SA·Parquet
Text

demo_data

1,000 examples from gpt4 en 1,000 examples from gpt4 zh 300 examples from toolcall en 300 examples from…

1K–10K·Apache-2.0
Text

evaluation-tables

!CAUTION This dataset will not be updated. It corresponds to the last available public snapshot of the data,…

1K–10K·CC-BY-SA·Parquet
MultimodalTabularText

Goodreads-Books

Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…

1M–10M·Custom / Research-only·CSV
MultimodalTextVideo

InsViE

InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction Citation If you find this work helpful, please consider…

1M–10M·CC-BY·CSV
MultimodalTextVideo

ProLongVid_data

Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…

1M–10M·Apache-2.0·JSON
Text

SWE-bench_Pro

Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…

<1K·Parquet
Text

TriviaQA

TriviaQA

100K–1M·Custom / Research-only·Parquet
Text

P3

P3

100M–1B·Apache-2.0·Parquet
MultimodalTabularText

common_corpus

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…

10K–100K·Parquet
Text

pile-uncopyrighted

Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The…

100M–1B·Custom / Research-only·JSON
Text

wikitext_document_level

Wikitext Document Level This is a modified version of that returns Wiki pages instead of Wiki text line-by-line.…

10K–100K·CC-BY-SA·Parquet
Advertisement