Skip to content
Advertisement

Visual Question Answering

339 results

ImageMultimodalText

Cambrian-Alignment

Cambrian-Alignment Dataset Please see paper & website for more information: Overview Cambrian-Alignment is an question-answering alignment dataset…

100K–1M·Apache-2.0·WebDataset
Image

Meissa-SFT

Meissa-SFT: Medical Agentic SFT Training Data This dataset contains the open-source portion of the SFT (Supervised Fine-Tuning) training…

10K–100K·Apache-2.0
ImageMultimodalText

FlowLearn

cr--- task categories: - visual-question-answering language: - en size categories: - 1K<n<10K References The original articles are maintained…

<1K·CC-BY-NC-SA·Images (folder)
ImageMultimodalText

XLRS-Bench-lite

🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…

1K–10K·CC-BY-NC-SA·Arrow
MultimodalTextVideo

heico-focus-vqa

HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper  • …

10K–100K·CC-BY-NC-SA·Parquet
Video

LLaVA-Video-large-swift

Dataset Card LLaVA-Video-medium-swift A subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.

<1K·Apache-2.0
ImageMultimodalText

EO-Data1.5M

🤖 EO-Data-1.5M A Large-Scale Interleaved Vision-Text-Action Dataset for Embodied AI The first large-scale interleaved embodied dataset emphasizing…

1M–10M·Apache-2.0·Parquet
Image

VLM-SubtleBench

VLM-SubtleBench VLM-SubtleBench: How Far Are VLMs from Human-Level Subtle Comparative Reasoning? The ability to distinguish subtle differences…

10K–100K·CC-BY-NC
ImageMultimodalText

Robo2VLM-1

Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets Abstract Vision-Language Models (VLMs) acquire…

100K–1M·Apache-2.0·Parquet
Text

HR-Bench

Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models 🌐Homepage 📖…

1K–10K·Custom / Research-only·Parquet
ImageMultimodalText

ULVR_v2_clean

ULVR v2 clean Universal Latent Visual Reasoning training data, cleaned. 8 categories (subsets); each has train + validation…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

CharXiv

CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs NeurIPS 2024 🏠Home (🚧Still in construction) 🤗Data 🥇Leaderboard…

1K–10K·CC-BY-SA·Parquet
Text

OmniCorpus-CC

🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text ⭐️ NOTE: Several parquet files…

100M–1B·CC-BY·Parquet
ImageMultimodalText

MAmmoTH-VL-Instruct-12M

MAmmoTH-VL-Instruct-12M 🏠 Homepage 🤖 MAmmoTH-VL-8B 💻 Code 📄 Arxiv 📕 PDF 🖥️ Demo Introduction Our simple yet scalable…

10M–100M·Apache-2.0·WebDataset
ImageMultimodalText

Zebra-CoT

Zebra‑CoT A diverse large-scale dataset for interleaved vision‑language reasoning traces. Dataset Description Zebra‑CoT is a diverse large‑scale…

100K–1M·CC-BY-NC·Parquet
Advertisement