Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
TimeGround-1M Synthetic English audio dataset for time-aware speech understanding, covering temporal localization, temporal description, and timed summaries. Data Filtering We use 14k hours of audio from YODAS2 English shards, selected from a 24k-hour source pool after language- and silence-ratio filtering. Synthetic annotations were generated for three time-grounded tasks, then filtered through LLM-based verification, deterministic validity checks, and… See the full description on the dataset page:
Source: Hugging Face Hub (ai-sage/TimeGround-1M). Metadata imported from the dataset’s Hub tags.