Skip to content
Advertisement

Audio Classification

81 results

AudioImageMultimodal

SoundingEarth

SoundingEarth SoundingEarth is a geo-referenced soundscape dataset that pairs Google Earth imagery with geotagged environmental audio recordings…

10K–100K·CC-BY·Parquet
AudioMultimodalText

bashkort_tts_dataset

Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…

10K–100K·CC-BY·Parquet
AudioMultimodalText

wuwa-voice-EN

wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total…

10K–100K·Parquet
AudioMultimodalTabular

YodasSpeakerPool

Use this dataset in conjuction with: YodasSpeakerPool YodasSpeakerPool is a curated, richly-annotated multi-speaker dataset featuring 7,600 unique…

1K–10K·CC-BY
AudioMultimodalText

SimbaBench_dataset

SibmaBench Data Release & Benchmarking To evaluate your model on SimbaBench across all supported tasks (ASR, TTS, and…

100K–1M·CC-BY·Parquet
AudioMultimodalText

C3T

C3T: Cross-modal Capabilities Conservation Test Dataset Description C3T (Cross-modal Capabilities Conservation Test) is a benchmark for assessing the…

100K–1M·CC-BY
Audio

emolia-thinking-balanced-buckets

Emolia-Thinking — Balanced Per-Dimension Bucket Subset A balanced, per-dimension bucket subset of VoiceNet/emolia-thinking, derived from that…

100K–1M·CC-BY
Advertisement