Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: Code & benchmarks: Paper: Files test.csv – original annotations and metadata with clip paths… See the full description on the dataset page:
Source: Hugging Face Hub (plnguyen2908/AV-SpeakerBench). Metadata imported from the dataset’s Hub tags.