Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Shofo Talking Head Dataset (English) A curated dataset of ~10,000 talking-head videos (~186 hours, mean ~67s/clip), filtered for clean single-speaker framing and paired with time-aligned transcripts. Built and released by Shofo. This dataset is short-form social video, designed to support modern avatar, lip-sync, dubbing, and TTS work. Every clip is curated by a multi-stage pipeline (face/framing analysis, on-screen-text detection, object-occlusion detection, voice/face matching… See the full description on the dataset page:
Source: Hugging Face Hub (Shofo/shofo-talking-head-en). Metadata imported from the dataset’s Hub tags.