Skip to content
Advertisement

Text

2,175 results

MultimodalTabularText

iris

Iris Species Dataset The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple…

<1K·CC0·CSV
3D / Point CloudMultimodalTabular

SynData

SynData 中文说明 Demo If the video cannot be displayed in your environment, open it directly: assets/syndata-demo.mp4 1. Overview…

100K–1M·CC-BY·Parquet
MultimodalTabularText

kinetics400

Kinetics-400 Video Dataset This dataset is derived from the Kinetics-400 dataset, which is released under the Creative Commons…

10K–100K·Parquet
MultimodalTabularText

car-bench-dataset

CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…

1M–10M·MIT·JSON
MultimodalTabularText

GlotCC-V1

Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is…

>1B·CC0
MultimodalTabularText

plurel

PluRel A large collection of 2000 synthetic relational databases (seeds plurel-0 … plurel-1999) generated for pretraining relational/tabular…

1K–10K·CC-BY-SA·Parquet
MultimodalTabularText

EuroWeb-2512

EuroWeb-2512 EuroWeb is a dataset of collecting multilingual web data from various sources. It was processed with standard…

>1B·Parquet
MultimodalTabularText

finemath

📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…

10M–100M·ODC-BY·Parquet
Advertisement