Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dataset Summary This is a 10K hours subset of English version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages – English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and… See the full description on the dataset page: eng 10k.
Source: Hugging Face Hub (parler-tts/mls_eng_10k). Metadata imported from the dataset’s Hub tags.