Skip to content
Advertisement
AudioMultimodalText

IndicVoices

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates 23 December 2025 We now have 11,200 hours of…

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates 23 December 2025 We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page:

Source: Hugging Face Hub (ai4bharat/IndicVoices). Metadata imported from the dataset’s Hub tags.

Advertisement