Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images,…
Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images, including charts, tables, documents, diagrams, etc. The synthetic data is generated by extending the CoSyn framework into multi-image settings, with Claude-sonnet-4-5 as the coding LLM to generate code that can be executed to render an image. Then, we use GPT-5 to generate question-answer pairs with code (without using the rendered… See the full description on the dataset page:
Source: Hugging Face Hub (allenai/Molmo2-SynMultiImageQA). Metadata imported from the dataset’s Hub tags.