Skip to content
Advertisement

Datasets

1,940 results

Text

VideoKR-Train

VideoKR-Train 📄 ArXiv  |  💻 Code  |  🤗 Collection About This repository contains the VideoKR training data presented…

100K–1M·Apache-2.0·JSON
ImageMultimodalText

Molmo2-SynMultiImageQA

Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images,…

100K–1M·ODC-BY·Parquet
ImageMultimodalText

ULVR-filtered

ULVR-filtered Filtered subset of RuoliuYang/ULVR v2 clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only…

100K–1M·Apache-2.0·Parquet
ImageMultimodalText

REVERSE

REVERSE Dataset Dataset for REVERSE (Reinforcing Evidence Verification and Search for Agentic Image Geolocation). This dataset supports training…

1M–10M·CC-BY·Parquet
ImageMultimodalText

GUIDE-dataset

GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation GUI Unbiasing via…

Apache-2.0
ImageMultimodalText

Kvasir-VQA-x1

Kvasir-VQA-x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir-VQA-x1 on GitHub Original Image…

100K–1M·CC-BY-NC·Parquet
ImageMultimodalText

ChartGalaxy

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation 🤗 Dataset 🖥️ Code 📄 Paper 🔥 News 2026.02…

1K–10K·CC-BY-NC·Parquet
ImageMultimodalText

RoboBench

RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain 📋 Overview RoboBench is a…

<1K·CC-BY
ImageMultimodalText

Cauldron-JA

Dataset Card for The Cauldron-JA Dataset description The Cauldron-JA is a Vision Language Model dataset that translates 'The…

1M–10M·CC-BY·Parquet
ImageMultimodalText

MMMU_Pro

MMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage 🏆 Leaderboard 🤗 Dataset 🤗 Paper 📖 arXiv…

1K–10K·Apache-2.0·Parquet
Advertisement