Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dhivehi Synthetic Voice and Speech Augmentation Dataset This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page:
Source: Hugging Face Hub (alakxender/dhivehi-audios-82-spk). Metadata imported from the dataset’s Hub tags.