Skip to content
Advertisement

10K–100K

748 results

Text

BoolQ

BoolQ

10K–100K·CC-BY-SA·Parquet
Text

C-Eval

C-Eval

10K–100K·CC-BY-NC-SA·Parquet
MultimodalTabularText

common_corpus

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…

10K–100K·Parquet
Text

wikitext_document_level

Wikitext Document Level This is a modified version of that returns Wiki pages instead of Wiki text line-by-line.…

10K–100K·CC-BY-SA·Parquet
Text

Alpaca

Alpaca

10K–100K·CC-BY-NC·Parquet
Text

SciQ

SciQ

10K–100K·Other·Parquet
Text

hendrycks_math

Dataset Summary MATH dataset from Citation Information @article{hendrycksmath2021, title={Measuring Mathematical Problem Solving With the MATH…

10K–100K·MIT·Parquet
Text

SWE-rebench

Dataset Summary SWE-rebench is a large-scale dataset designed to support training and evaluation of LLM-based software engineering (SWE)…

10K–100K·CC-BY·Parquet
Text

SQuAD

SQuAD

10K–100K·CC-BY-SA·Parquet
ImageMultimodalText

COCO-Caption

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10K–100K·Parquet
Advertisement