Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page:
Source: Hugging Face Hub (ghananlpcommunity/kasem-speech-text-parallel). Metadata imported from the dataset’s Hub tags.