Skip to content
Advertisement

Datasets

3,216 results

ImageMultimodalText

ship-dataset

ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on…

10K–100K·CC-BY·JSON
ImageMultimodalText

Obshazard-bench

ObsCrisis-Bench A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks. Dataset Description…

<1K·MIT
Video

stream-data

Streaming Video Dataset Description A consolidated collection of video datasets for streaming video understanding research, including temporal…

100K–1M·CC-BY
ImageMultimodalText

CoSyn-400K

CoSyn-400k CoSyn-400k is a collection of synthetic question-answer pairs about very diverse range of computer-generated images. The data…

100K–1M·ODC-BY·Parquet
MultimodalTextVideo

UrbanVideo-Bench

ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…

1K–10K·MIT·Parquet
ImageMultimodalText

Cambrian-Alignment

Cambrian-Alignment Dataset Please see paper & website for more information: Overview Cambrian-Alignment is an question-answering alignment dataset…

100K–1M·Apache-2.0·WebDataset
Image

Meissa-SFT

Meissa-SFT: Medical Agentic SFT Training Data This dataset contains the open-source portion of the SFT (Supervised Fine-Tuning) training…

10K–100K·Apache-2.0
ImageMultimodalText

FlowLearn

cr--- task categories: - visual-question-answering language: - en size categories: - 1K<n<10K References The original articles are maintained…

<1K·CC-BY-NC-SA·Images (folder)
ImageMultimodalText

XLRS-Bench-lite

🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…

1K–10K·CC-BY-NC-SA·Arrow
MultimodalTextVideo

heico-focus-vqa

HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper  • …

10K–100K·CC-BY-NC-SA·Parquet
Video

LLaVA-Video-large-swift

Dataset Card LLaVA-Video-medium-swift A subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.

<1K·Apache-2.0
ImageMultimodalText

EO-Data1.5M

🤖 EO-Data-1.5M A Large-Scale Interleaved Vision-Text-Action Dataset for Embodied AI The first large-scale interleaved embodied dataset emphasizing…

1M–10M·Apache-2.0·Parquet
Advertisement