Skip to content
Advertisement

Text

2,175 results

MultimodalTabularText

hub-stats

Changelog NEW Changes March 11th 2026 Added new split: arxiv papers, sourced from the Hugging Face /api/papers endpoint…

1M–10M·Apache-2.0·Parquet
MultimodalTabularText

MMLU-ProX

MMLU-ProX MMLU-ProX is a multilingual benchmark that builds upon MMLU-Pro, extending to 29 typologically diverse languages, designed to…

100K–1M·MIT·Parquet
MultimodalTabularText

OceanTACO

Dataset Card: OceanTACO Dataset Summary This dataset is a multi-source collection of global ocean sea surface measurements, integrating…

10K–100K·CC-BY
MultimodalTabularText

demo1

Dataset Card for Demo1 Dataset Summary This is a demo dataset. It consists in two files data/train.csv and…

<1K·CSV
MultimodalTabularText

betty-dota2

Betty Dota 2 — Decision Context Dataset Overview 9,385 professional Dota 2 matches parsed from replay files (.dem)…

>1B·MIT·Parquet
MultimodalTabularText

redelex

CTU relational datasets (redelex) Relational databases from the CTU Prague Relational Learning Repository (a.k.a. the CTU relational repository),…

<1K·CC-BY-SA·Parquet
MultimodalTabularText

qm9

Dataset Card for "QM9" QM9 dataset from Ruddigkeit et al., 2012; Ramakrishnan et al., 2014. Original data downloaded…

100K–1M·Parquet
MultimodalTabularText

cc100-documents

cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…

100M–1B·Parquet
MultimodalTabularText

liars-bench

Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.…

10K–100K·Custom / Research-only·Parquet
MultimodalTabularText

codeforces-cots

Dataset Card for CodeForces-CoTs Dataset description CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive…

100K–1M·CC-BY·Parquet
Advertisement