Skip to content
Advertisement

Open

2,922 results

MultimodalTabularText

plurel

PluRel A large collection of 2000 synthetic relational databases (seeds plurel-0 … plurel-1999) generated for pretraining relational/tabular…

1K–10K·CC-BY-SA·Parquet
MultimodalTabularText

EuroWeb-2512

EuroWeb-2512 EuroWeb is a dataset of collecting multilingual web data from various sources. It was processed with standard…

>1B·Parquet
MultimodalTabularText

finemath

📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…

10M–100M·ODC-BY·Parquet
MultimodalTabularText

Polymarket_data

Polymarket Data Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze A comprehensive dataset of 1.9 billion trading…

>1B
MultimodalTabularText

molmobot-data

MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:…

100K–1M·ODC-BY·Parquet
Video

FluxBisimData

FluxBisimData FluxBisimData is a simulated manipulation dataset collected by FluxBisim and used for the model evaluation of FluxBisim.…

1K–10K·Apache-2.0
Video

X-WAM-RoboCasa

X-WAM Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Dataset Summary This is the RoboCasa…

1K–10K·Apache-2.0
MultimodalTextVideo

VSI-Bench

Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased !IMPORTANT Nov. 7, 2025 UPDATE: This Dataset has been updated to include…

10K–100K·Apache-2.0·Parquet
Video

THQA-NTIRE

🚀 CVPR NTIRE 2025 - XGC Quality Assessment - Track 3: Talking Head (THQA-NTIRE) "Who is a Better…

10K–100K·MIT
Video

ActivityNet

Description Dataset V1-2 v1-2 train.tar.gz and v1-2 val.tar.gz Data (train and val set) associated with ActivityNet release 1.2…

<1K
Video

AVUTBenchmark

Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text…

1K–10K
Advertisement