Skip to content
Advertisement
ImageMultimodalText

Molmo2-SynMultiImageQA

Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images,…

Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images, including charts, tables, documents, diagrams, etc. The synthetic data is generated by extending the CoSyn framework into multi-image settings, with Claude-sonnet-4-5 as the coding LLM to generate code that can be executed to render an image. Then, we use GPT-5 to generate question-answer pairs with code (without using the rendered… See the full description on the dataset page:

Source: Hugging Face Hub (allenai/Molmo2-SynMultiImageQA). Metadata imported from the dataset’s Hub tags.

Advertisement