llava-en-zh-300k
This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…
2,147 results
This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…
MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal…
MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…
PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…
Nornikel Metallurgy VL Dataset
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…
Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…
Indian Competitive Exams (JEE/NEET) LLM Benchmark
Video-opd-Dataset A video temporal grounding dataset with 2,500 samples sourced from TimeLens-100K. Dataset Description This dataset contains video…
ChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…
Molmo2 PointArena SFT Data (LAION + Molmo-7B-D)
Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…
Optimism Bias World Model Benchmark — Training Data Training data for VLM-as-judge models evaluating world model predictions. Structure…
MVTamperBench Dataset Overview MVTamperBenchEnd is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
ICML 2026 Yiming Zhang1 , Jiacheng Chen1 , Jiaqi Tan1, Yongsen Mao2, Wenhu Chen3, Angel X. Chang1,4 1…
AVQA (Audio-Visual QA) — Videos + Annotations