Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
PVIT-3M The paper titled “Personalized Visual Instruction Tuning” introduces a novel dataset called PVIT-3M. This dataset is specifically designed for tuning MLLMs in the context of personalized visual instruction tasks. The dataset consists of 3 million image-text pairs that aim to improve MLLMs’ abilities to generate responses based on personalized visual inputs, making them more tailored and adaptable to individual user needs and preferences. Here’s the PVIT-3M statistics:… See the full description on the dataset page:
Source: Hugging Face Hub (Sterzhang/PVIT-3M). Metadata imported from the dataset’s Hub tags.