Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
YTSeg: A Benchmark for Audio Chaptering and Video Transcript Segmentation We present YTSeg, a topically and structurally diverse benchmark for the audio chaptering and transcript segmentation task based on YouTube videos. The dataset comprises 19,299 videos from 393 channels, amounting to 6,533 content hours. The topics are wide-ranging, covering domains such as science, lifestyle, politics, health, economy, and technology. The videos are from various types of content formats… See the full description on the dataset page:
Source: Hugging Face Hub (retkowski/ytseg). Metadata imported from the dataset’s Hub tags.