Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Yoruba Speech-Text Parallel Dataset Dataset Description This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in Nigeria and other West African countries. The dataset consists of audio recordings paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks. Dataset Summary Language: Yoruba – yo Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page:
Source: Hugging Face Hub (michsethowusu/yoruba-speech-text-parallel). Metadata imported from the dataset’s Hub tags.