Skip to content
Advertisement

Automatic Speech Recognition

267 results

AudioMultimodalText

UltraVoice

UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack…

100K–1M·MIT·JSON
AudioMultimodalText

MiscSpeech-ja

MiscSpeech-ja This dataset comprises audio and corresponding transcripts collected from a diverse range of YouTube videos and Podcasts.

10K–100K·CC-BY-SA·Parquet
AudioMultimodalText

vctk

VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero…

10K–100K·CC-BY·Parquet
AudioMultimodalText

Emilia-NV

NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at…

100K–1M·CC-BY-NC-SA·WebDataset
Audio

hawrami-kurdish-raw-audio

Hawrami Raw Audio Collection Overview This repository contains approximately 500 hours of Hawrami Kurdish raw speech collected from…

<1K·Audio (folder)
MultimodalTabularText

UniST

UniST This dataset contains UniST codec-token training data exported from local metadata and codec results. We train UniSS…

10M–100M·CC-BY-NC
AudioMultimodalText

Magpie-Speech-Orpheus-125k

Magpie-Speech-Orpheus-125k A ~125k-sample synthetic speech dataset generated by applying the Magpie instruction-synthesis approach to the Orpheus-TTS…

100K–1M·Other·Parquet
Audio

southern-kurdish-raw-audio

Southern Kurdish Raw Audio Collection Overview This repository contains approximately 170 hours of Southern Kurdish (SDH) raw speech…

<1K·Audio (folder)
Advertisement