Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
GRIT: Large-Scale Training Corpus of Grounded Image-Text Pairs Dataset Summary We introduce GRIT, a large-scale dataset of Grounded Image-Text pairs, which is created based on image-text pairs from COYO-700M and LAION-2B. We construct a pipeline to extract and link text spans (i.e., noun phrases, and referring expressions) in the caption to their corresponding image regions. More details can be found in the paper. Supported Tasks During the construction, we… See the full description on the dataset page:
Source: Hugging Face Hub (zzliang/GRIT). Metadata imported from the dataset’s Hub tags.