Skip to content
Advertisement
Video

RoVid-X

Rethinking Video Generation Model for the Embodied World If you like our project, please give us a star ⭐ on GitHub for the latest update. Key…

Rethinking Video Generation Model for the Embodied World If you like our project, please give us a star ⭐ on GitHub for the latest update. Key features 4M robotic video clips(10K+ hours) for large-scale video generation training. 1300+ fine-grained robotic skills, covering diverse actions and task primitives. Multi-modal physical annotations, including RGB, depth, and optical flow. Multi-robot and multi-task diversity… See the full description on the dataset page:

Source: Hugging Face Hub (DAGroup-PKU/RoVid-X). Metadata imported from the dataset’s Hub tags.

Advertisement