Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used…
LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used in mid-training. Dataset Composition Subset Format Description mid training video/60s rest/ WebDataset (.tar) 10,809 shards of ~60s video clips mid training video/caption v0/split 30s.jsonl JSONL Captions for 30-second video clips mid training video/caption v0/split 60s.jsonl JSONL Captions for… See the full description on the dataset page:
Source: Hugging Face Hub (mvp-lab/LLaVA-OneVision-2-Data). Metadata imported from the dataset’s Hub tags.