Skip to content
Advertisement
MultimodalTabularText

mumospee_emilia

Dataset Summary This dataset is a modified version of the Emilia corpus, converted into parquet format to facilitate optimized I/O operations in…

Dataset Summary This dataset is a modified version of the Emilia corpus, converted into parquet format to facilitate optimized I/O operations in high-performance and distributed computing environments. The Emilia dataset is a comprehensive, multilingual dataset with the following features: containing over 101k hours of speech data; covering six different languages: English (En), Chinese (Zh), German (De), French (Fr), Japanese (Ja), and Korean (Ko); containing diverse speech… See the full description on the dataset page: emilia.

Source: Hugging Face Hub (meetween/mumospee_emilia). Metadata imported from the dataset’s Hub tags.

Advertisement