Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Tamazight-Arabic Speech Translation Dataset
Tamazight-Arabic Speech Recognition Dataset This is the Tamazight-NLP organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset. This dataset contains ~15.5 hours of Tamazight (Tachelhit dialect) speech paired with Arabic transcriptions, designed for automatic speech recognition (ASR) and speech-to-text translation tasks. Dataset Details Total Examples: 20,344 audio segments Training Set: 18,309 examples (~8.9GB) Test Set: 2,035 examples (~992MB)… See the full description on the dataset page:
Source: Hugging Face Hub (Tamazight-NLP/Tamazight-Speech-to-Arabic-Text). Metadata imported from the dataset’s Hub tags.