Skip to content
Advertisement

Automatic Speech Recognition

267 results

AudioMultimodalText

Habibi

A systematic and standardized benchmark for the multi-dialect Arabic zero-shot TTS task Paper: "Habibi: Laying the Open-Source Foundation…

10K–100K·Apache-2.0·Parquet
Audio

Urdu-Munch-Lina

Urdu-Munch-Lina Processed version of zuhri025/Urdu-Munch with LinaCodec encoding. Dataset Structure This dataset contains 2 batches of audio data…

10K–100K·MIT
AudioMultimodalText

bashkort_tts_dataset

Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…

10K–100K·CC-BY·Parquet
AudioMultimodalText

live-atc-europe

Live ATC Europe — audio + transcriptions Enregistrements live d'air traffic control (ATC) européen, capturés depuis LiveATC.net, segmentés…

100K–1M·Custom / Research-only·Parquet
Advertisement