Skip to content
Advertisement

Gated

293 results

Featured
MultimodalVideo

Ego4D

3,670 hours of first-person video across hundreds of scenarios.

1M–10M·Custom / Research-only·Video (mp4)
Image

GeoHop

GeoHop — Multi-Hop Remote-Sensing VQA (benchmarks + images + weights) Data/weights mirror for Da1daidaidai/GeoHop (code, training, eval). This…

1K–10K·CC-BY·Images (folder)
Multimodal

Changen2-S9-27k

Dataset Card for Changen2-S9-27k Changen2-S9-27k (an urban land-use/landcover change dataset with 27k pairs and 38 change types), 0.25-0.5m…

10K–100K·CC-BY-NC-SA
Image

EarthVLSet

EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework by Junjue Wang, Yanfei Zhong, Zihang Chen, Zhuo Zheng,…

>1B·CC-BY-NC-ND
Text

nmt-parallel-corpus

Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned…

>1B·CC-BY-NC·Parquet
Text

iSign

iSign

100K–1M·CC-BY-NC-SA·CSV
MultimodalTextVideo

lapchole-focus-vqa

LapChole-FOCUS-VQA A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. 💻 Code  •  🏆…

10K–100K·Custom / Research-only·Parquet
MultimodalTabularText

CG-Bench

CG-Bench Project Website: Repository: (includes running code) Summary We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question…

10K–100K·MIT·JSON
Video

stream-data

Streaming Video Dataset Description A consolidated collection of video datasets for streaming video understanding research, including temporal…

100K–1M·CC-BY
MultimodalTextVideo

heico-focus-vqa

HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper  • …

10K–100K·CC-BY-NC-SA·Parquet
Advertisement