Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
AVQA (Audio-Visual QA) — Videos + Annotations
AVQA — Audio-Visual Question Answering (videos + annotations) A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life audio-visual question answering over short in-the-wild clips. The original release ships only the QA annotations and expects users to collect the source videos from VGGSound themselves. This repository bundles the source video clips together with the official train/val annotations, so the dataset is usable without any YouTube scraping.… See the full description on the dataset page:
Source: Hugging Face Hub (juyil/AVQA-videos). Metadata imported from the dataset’s Hub tags.