Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Multimodal Video QA Dataset This dataset contains challenging video question-answering tasks that require understanding both visual and audio…
Multimodal Video QA Dataset This dataset contains challenging video question-answering tasks that require understanding both visual and audio information across entire video timelines. Dataset Statistics Total Videos: 4,140 Total Size: 118.08 GB Dataset Structure The dataset is split into multiple parts (each ≤2GB): Part 1: 75 videos (1.97 GB) – videos part001.zip Part 2: 49 videos (1.97 GB) – videos part002.zip Part 3: 55 videos (1.91 GB) -… See the full description on the dataset page:
Source: Hugging Face Hub (ngqtrung/full-modality-bench). Metadata imported from the dataset’s Hub tags.