AVQA (Audio-Visual QA) — Videos + Annotations
AVQA (Audio-Visual QA) — Videos + Annotations
911 results
AVQA (Audio-Visual QA) — Videos + Annotations
Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…
MVTamperBench Dataset Overview MVTamperBenchStart is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
CG-Bench Project Website: Repository: (includes running code) Summary We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question…
Dataset Info: SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering ISBI 2021 oral Project Page: click…
ESpeech datasets annotate by Balalaika
Speech DAC Tokens (3 Codebooks)
Dutch TTS Dataset - Complete Labeled A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of…
Dataset Card for Filtred and annotated CML TTS This dataset is an annotated and filtred version of a…
Dataset Card for Annotations of non English MLS This dataset consists in annotations of a the Non English…
Risale-i Nur Sesli Külliyat — Risale-i Nur Audio Corpus
MintTTS Pre-tokenized Audio (somu9/hindi-hq)
UniST This dataset contains UniST codec-token training data exported from local metadata and codec results. We train UniSS…
VoxCPM Ghana — Precomputed AudioVAE Latents
Use this dataset in conjuction with: YodasSpeakerPool YodasSpeakerPool is a curated, richly-annotated multi-speaker dataset featuring 7,600 unique…
OpenSTT annotate by Balalaika