Skip to content
Advertisement
ImageMultimodalTextVideo

ShareGPT4Video Captions Dataset Card

ShareGPT4Video Captions Dataset Card

ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video 40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page:

Source: Hugging Face Hub (ShareGPT4Video/ShareGPT4Video). Metadata imported from the dataset’s Hub tags.

Advertisement