Skip to content
Advertisement
Video

X-WAM-RoboCasa

X-WAM Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Dataset Summary This is the RoboCasa fine-tuning dataset used to…

X-WAM Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Dataset Summary This is the RoboCasa fine-tuning dataset used to train the X-WAM unified 4D World Action Model. It packages single-arm kitchen manipulation demonstrations into a unified multi-view RGB-D video + low-dimensional state/action format, where each episode provides synchronized RGB videos, depth videos, end-effector proprioception, actions, and language… See the full description on the dataset page:

Source: Hugging Face Hub (sharinka0715/X-WAM-RoboCasa). Metadata imported from the dataset’s Hub tags.

Advertisement