Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
WorldSpeech A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language-region variants, drawn from national parliaments, public broadcasters, public-domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human-provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA-SNR estimate, and four DNSMOS-P.835 quality scores. Dataset Overview… See the full description on the dataset page:
Source: Hugging Face Hub (disco-eth/WorldSpeech). Metadata imported from the dataset’s Hub tags.