Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Syrian Postcast Arabic Speech Dataset Dataset Summary The Syrian Postcast Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 116 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page:
Source: Hugging Face Hub (oddadmix/arabic-audio-collection-syrian-podcast). Metadata imported from the dataset’s Hub tags.