Skip to content
Advertisement
Video

full-modality-bench

Multimodal Video QA Dataset This dataset contains challenging video question-answering tasks that require understanding both visual and audio…

Multimodal Video QA Dataset This dataset contains challenging video question-answering tasks that require understanding both visual and audio information across entire video timelines. Dataset Statistics Total Videos: 4,140 Total Size: 118.08 GB Dataset Structure The dataset is split into multiple parts (each ≤2GB): Part 1: 75 videos (1.97 GB) – videos part001.zip Part 2: 49 videos (1.97 GB) – videos part002.zip Part 3: 55 videos (1.91 GB) -… See the full description on the dataset page:

Source: Hugging Face Hub (ngqtrung/full-modality-bench). Metadata imported from the dataset’s Hub tags.

Advertisement