Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Crowd Recital Yiddish - Source Dataset
About This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project. Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read. Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page:
Source: Hugging Face Hub (ivrit-ai/crowd-recital-yi). Metadata imported from the dataset’s Hub tags.