Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dataset Card for Voxpopuli Dataset Summary VoxPopuli is a large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. The raw data is collected from 2009-2020 European Parliament event recordings. We acknowledge the European Parliament for creating and sharing these materials. This implementation contains transcribed speech data for 18 languages. It also contains 29 hours of transcribed speech data of non-native English… See the full description on the dataset page:
Source: Hugging Face Hub (facebook/voxpopuli). Metadata imported from the dataset’s Hub tags.