Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
ShareGPT4Video 4.8M Dataset Card Dataset details Dataset type: ShareGPT4Video Captions 4.8M is a set of GPT4-Vision-powered multi-modal captions data of videos. It is constructed to enhance modality alignment and fine-grained visual concept perception in Large Video-Language Models (LVLMs) and Text-to-Video Models (T2VMs). This advancement aims to bring LVLMs and T2VMs towards the capabilities of GPT4V and Sora. sharegpt4video 40k.jsonl is generated by GPT4-Vision… See the full description on the dataset page:
Source: Hugging Face Hub (ShareGPT4Video/ShareGPT4Video). Metadata imported from the dataset’s Hub tags.