Skip to content
Advertisement

Image

1,418 results

ImageMultimodalText

llava-en-zh-300k

This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…

100K–1M·Apache-2.0·Parquet
ImageMultimodalTabular

MVBench

MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…

1K–10K·MIT·JSON
ImageMultimodalText

pixmo-docs

PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…

100K–1M·ODC-BY·Parquet
ImageMultimodalText

PhyX

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…

10K–100K·MIT·Parquet
ImageMultimodalText

VLAA-Thinking

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…

<1K·Apache-2.0·Images (folder)
Image

PanoEnv

CVPR 2026 Highlight - PanoEnv-QA: A Large-Scale Geometry-Grounded Panoramic VQA Benchmark for 3D Spatial Intelligence 📖 Overview PanoEnv-QA…

10K–100K·CC-BY·Images (folder)
ImageMultimodalText

ChartVerse-SFT-600K

ChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…

100K–1M·Apache-2.0·Parquet
Image

SandThink

SandThink Dataset (v1.0) SandThink 是一个专为具身智能 (Embodied AI) 任务设计的大规模指令微调与偏好对齐数据集。该数据集通过结构化的 Chain-of-Thought (CoT) 推理过程,显著提升了 Vision-Language-Action…

<1K·MIT·Images (folder)
Advertisement