Skip to content
Advertisement
MultimodalTabularText

Wan2.2-Syn-121x704x1280_32k

FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by…

FastVideo Synthetic Wan2.2 720P dataset FastVideo Team Paper Github Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at emph{both} training and inference. In VSA, a… See the full description on the dataset page: 32k.

Source: Hugging Face Hub (Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k). Metadata imported from the dataset’s Hub tags.

Advertisement