Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
DoclingMatix DoclingMatix is a large-scale, multimodal dataset designed for training vision-language models in the domain of document intelligence. It was created specifically for training the SmolDocling model, an ultra-compact model for end-to-end document conversion. The dataset is constructed by augmenting Hugging Face’s Docmatix. Each sample in Docmatix, which consists of a document image and a few questions and answers about it, has been transformed. The text field is now… See the full description on the dataset page:
Source: Hugging Face Hub (HuggingFaceM4/DoclingMatix). Metadata imported from the dataset’s Hub tags.