Skip to content
Advertisement
AudioMultimodalTextVideo

AV-SpeakerBench

AV-SpeakerBench

AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: Code & benchmarks: Paper: Files test.csv – original annotations and metadata with clip paths… See the full description on the dataset page:

Source: Hugging Face Hub (plnguyen2908/AV-SpeakerBench). Metadata imported from the dataset’s Hub tags.

Advertisement