Skip to content
Advertisement

Dataset Card for English MLS Dataset Summary This is a streamable version of the English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages – English, German, Dutch, Spanish, French, Italian, Portuguese… See the full description on the dataset page:

Source: Hugging Face Hub (ntt123/mls-eng-128kb). Metadata imported from the dataset’s Hub tags.

Advertisement