Skip to content
Advertisement
AudioMultimodalText

MultiLingual LibriSpeech

MultiLingual LibriSpeech

Dataset Card for MultiLingual LibriSpeech Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages – English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: librispeech.

Source: Hugging Face Hub (facebook/multilingual_librispeech). Metadata imported from the dataset’s Hub tags.

Advertisement