Skip to content
Advertisement
MultimodalTextVideo

Vript_Chinese

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 44.7K annotated high-resolution…

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 44.7K annotated high-resolution videos (~293k clips) in Chinese. The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close-up, etc), and how the camera moves (panning… See the full description on the dataset page: Chinese.

Source: Hugging Face Hub (Mutonix/Vript_Chinese). Metadata imported from the dataset’s Hub tags.

Advertisement