Skip to content
Advertisement

English

1,233 results

3D / Point CloudMultimodalTabular

SynData

SynData 中文说明 Demo If the video cannot be displayed in your environment, open it directly: assets/syndata-demo.mp4 1. Overview…

100K–1M·CC-BY·Parquet
MultimodalTabularText

car-bench-dataset

CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…

1M–10M·MIT·JSON
MultimodalTabularText

molmobot-data

MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:…

100K–1M·ODC-BY·Parquet
Video

FluxBisimData

FluxBisimData FluxBisimData is a simulated manipulation dataset collected by FluxBisim and used for the model evaluation of FluxBisim.…

1K–10K·Apache-2.0
MultimodalTextVideo

VSI-Bench

Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased !IMPORTANT Nov. 7, 2025 UPDATE: This Dataset has been updated to include…

10K–100K·Apache-2.0·Parquet
Video

AVUTBenchmark

Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text…

1K–10K
Video

hoigen-filtered-videos

HOIGen Filtered Videos Dataset This dataset contains 28562 filtered videos from the HOIGen-1M dataset based on the allowlist.…

1K–10K·MIT
Advertisement