Skip to content
Advertisement

Parquet

1,501 results

MultimodalTabularText

fake_news

TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: annotations creators: - no-annotation…

10K–100K·Parquet
ImageMultimodalTabular

filtered-wit

Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image…

1M–10M·Parquet
MultimodalTabularText

wildguardmix

Dataset Card for WildGuardMix Disclaimer: The data includes examples that might be disturbing, harmful or upsetting. It includes…

10K–100K·ODC-BY·Parquet
MultimodalTabularText

SurvHTE-Bench

SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis Paper: ICLR 2026 — SurvHTE-Bench: A Benchmark…

1M–10M·CC-BY·Parquet
MultimodalTabularText

ScaleEdit-12M

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework       📌 Overview The largest open-source…

10M–100M·CC-BY-NC-SA·Parquet
MultimodalTabularText

codeforces

Dataset Card for CodeForces Dataset description CodeForces is one of the most popular websites among competitive programmers, hosting…

10K–100K·CC-BY·Parquet
MultimodalTabularText

leaderboard-dataset

Arena Leaderboard Dataset Historical snapshots of the Arena leaderboard. Usage from datasets import load dataset Load all historical…

1M–10M·CC-BY·Parquet
ImageMultimodalTabular

datacomp_1b

DataComp-1B This repository contains metadata files for DataComp-1B. For details on how to use the metadata, please visit…

>1B·CC-BY·Parquet
MultimodalTabularText

fineweb-2-edu-japanese

🍷 FineWeb2 Edu Japanese: High-Quality Educational Japanese Dataset This dataset consists of 120 million texts (approximately 89.3B tokens)…

100M–1B·ODC-BY·Parquet
Advertisement