Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dataset Description, Collection, and Source The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses. In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page:
Source: Hugging Face Hub (FBK-MT/mosel). Metadata imported from the dataset’s Hub tags.