Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for…
Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B answer tokens, designed for training state-of-the-art open Vision-Language-Models. More detail can be found in the blog post: Load the data from datasets import load dataset, get dataset config names Get all subset names and load the first one available subsets =… See the full description on the dataset page:
Source: Hugging Face Hub (HuggingFaceM4/FineVision). Metadata imported from the dataset’s Hub tags.