Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
LargeScaleASR: 25,000 hours of transcribed and heterogeneous English speech recognition data for research and commercial use. The full details are available in the paper. Made of 6 subsets: large contains 25,000 hours of read / spontaneous and clean / noisy transcribed speech. medium contains 2,500 hours of read / spontaneous and clean / noisy transcribed speech. small contains 250 hours of read / spontaneous and clean / noisy transcribed speech. clean contains 13,000 hours of read… See the full description on the dataset page:
Source: Hugging Face Hub (speechbrain/LoquaciousSet). Metadata imported from the dataset’s Hub tags.