Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text…
Audio-centric Video Understanding Benchmark (AVUT) This dataset is presented in the paper Audio-centric Video Understanding Benchmark without Text Shortcut. Code Repository: Paper: Introduction The Audio-centric Video Understanding Benchmark (AVUT) aims to evaluate the video comprehension capabilities of multimodal Large Language Models (LLMs), with a particular focus on auditory information. Audio… See the full description on the dataset page:
Source: Hugging Face Hub (tsinghua-ee/AVUTBenchmark). Metadata imported from the dataset’s Hub tags.