Ego4D
3,670 hours of first-person video across hundreds of scenarios.
293 results
3,670 hours of first-person video across hundreds of scenarios.
1,000 driving scenes with 360° camera, LiDAR and radar.
Dataset Card for Changen2-S9-27k Changen2-S9-27k (an urban land-use/landcover change dataset with 27k pairs and 38 change types), 0.25-0.5m…
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework by Junjue Wang, Yanfei Zhong, Zihang Chen, Zhuo Zheng,…
KYNMAW — Khasi Bhashini Multi-Task Dataset
Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned…
Turkish Academic Theses Abstracts (TR/EN)
African Languages Lab Multi-Open
DCVLM-Baseline (200B tokens)
Dataset Card for OpenDocVQA This is a training and evaluation corpus data file for VDocRAG, a new RAG…
LapChole-FOCUS-VQA A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. 💻 Code • 🏆…
CG-Bench Project Website: Repository: (includes running code) Summary We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question…
Streaming Video Dataset Description A consolidated collection of video datasets for streaming video understanding research, including temporal…
HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper • …