CC-BY
672 results
RoboBench
RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain 📋 Overview RoboBench is a…
Cauldron-JA
Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates 'The…
IndicTTS_Bengali
Bengali Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Bengali…
Bambara-ASR-All Audio Dataset
Bambara-ASR-All Audio Dataset
InfoRe Technology public dataset №1
InfoRe Technology public dataset №1
Lahgtna Levantine TTS — Synthetic Levantine Arabic & Code-Switching
Lahgtna Levantine TTS — Synthetic Levantine Arabic & Code-Switching
Barranquenho IPA Pronunciation Dictionary
Barranquenho IPA Pronunciation Dictionary
Common Voice 13 French (Phonemized & Curated)
Common Voice 13 French (Phonemized & Curated)
bashkort_tts_dataset
Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…
Nagatoro Hayase Voice Dataset (140 Clips)
Nagatoro Hayase Voice Dataset (140 Clips)
libritts_p_dataset_20250821_095157
Contribution This dataset is a processed version of the original LibriTTS-P dataset, optimized for use on the Hugging…
WorldAudioNaturalConversations Sample Dataset
WorldAudioNaturalConversations Sample Dataset
Speech DAC Tokens (3 Codebooks)
Speech DAC Tokens (3 Codebooks)
libritts_r
Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…
dutch-tts-labeled-complete
Dutch TTS Dataset - Complete Labeled A comprehensive Dutch text-to-speech dataset with 596,508 audio samples totaling 234GB of…