Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: audio alignments.
Source: Hugging Face Hub (takuM23/multilingual_audio_alignments). Metadata imported from the dataset’s Hub tags.