Skip to content
Advertisement
AudioMultimodalText

Uzbek YouTube Speech Dataset

Uzbek YouTube Speech Dataset

Dataset Description A speech dataset of Uzbek language audio clips sourced from YouTube videos. Audio segments were extracted, separated by speaker using vocal isolation, and transcribed using Google’s Gemini 2.0 Flash model. Speaker identities were clustered using ECAPA-TDNN embeddings. Use Cases Automatic Speech Recognition (ASR) for Uzbek Text-to-Speech (TTS) synthesis for Uzbek Fine-tuning speech models on Uzbek language data (e.g., Qwen3-TTS) Speaker-conditioned TTS… See the full description on the dataset page: transcriptions.

Source: Hugging Face Hub (openbank-uz/youtube_transcriptions). Metadata imported from the dataset’s Hub tags.

Advertisement