Skip to content
Advertisement

Automatic Speech Recognition

267 results

Audio

Marco_Longspeech

Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks…

10K–100K·Apache-2.0
AudioMultimodalText

X-Voice-Dataset-Train

X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech…

10M–100M·Custom / Research-only·WebDataset
AudioMultimodalText

EuroSpeech

EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…

10M–100M·Custom / Research-only·Parquet
AudioMultimodalText

yodas2_sidon

YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis…

1M–10M·CC-BY·WebDataset
AudioMultimodalText

coraal

Corpus of Regional African American Language (CORAAL) Dataset link: Hugging Face Hub preparation scripts: coraal Citation Kendall, Tyler…

<1K·CC-BY-NC-SA·Parquet
Advertisement