Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
MASC Arabic Dataset Card Dataset Summary MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels. The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page:
Source: Hugging Face Hub (MohamedRashad/MASC-Arabic). Metadata imported from the dataset’s Hub tags.