Image
1,418 results
llava-en-zh-300k
This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…
FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal…
MVBench
MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…
pixmo-docs
PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…
Visual Intelligence Leaderboard
Visual Intelligence Leaderboard
Nornikel Metallurgy VL Dataset
Nornikel Metallurgy VL Dataset
PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
VLAA-Thinking
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…
Indian Competitive Exams (JEE/NEET) LLM Benchmark
Indian Competitive Exams (JEE/NEET) LLM Benchmark
ChartVerse-SFT-600K
ChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…
Molmo2 PointArena SFT Data (LAION + Molmo-7B-D)
Molmo2 PointArena SFT Data (LAION + Molmo-7B-D)