Datasets
3,216 results
FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal…
MVBench
MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…
pixmo-docs
PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…
Visual Intelligence Leaderboard
Visual Intelligence Leaderboard
Nornikel Metallurgy VL Dataset
Nornikel Metallurgy VL Dataset
PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
OpenDocVQA
Dataset Card for OpenDocVQA This is a training and evaluation QA data file for VDocRAG, a new RAG…
VLAA-Thinking
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…
VideoMarathon
Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…
Indian Competitive Exams (JEE/NEET) LLM Benchmark
Indian Competitive Exams (JEE/NEET) LLM Benchmark
synthvision-validated-qwen-by-kimi
synthvision-validated-qwen-by-kimi Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate) Records: 55,359 About Cross-validated subset from…
Video-opd-Dataset
Video-opd-Dataset A video temporal grounding dataset with 2,500 samples sourced from TimeLens-100K. Dataset Description This dataset contains video…
ChartVerse-SFT-600K
ChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…