Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by…
FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at emph{both} training and inference. In VSA, a… See the full description on the dataset page: 32k.
Source: Hugging Face Hub (Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k). Metadata imported from the dataset’s Hub tags.