Skip to content
Advertisement
AudioMultimodalText

dv_syn_audios

dv syn audios

Dhivehi Synthetic Voice and Speech Augmentation Dataset This dataset is a multi-speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice-cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice-representation research in low-resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre… See the full description on the dataset page:

Source: Hugging Face Hub (alakxender/dhivehi-audios-82-spk). Metadata imported from the dataset’s Hub tags.

Advertisement