Skip to content
Advertisement

Datasets

1,940 results

ImageMultimodalText

FineBench

Dataset Card for FineBench FineBench is a large-scale, multiple-choice Video Question Answering (VQA) dataset designed specifically to evaluate…

100K–1M·MIT·JSON
MultimodalTextVideo

VideoKR-Eval

VideoKR-Eval 📄 ArXiv  |  💻 Code  |  🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…

1K–10K·Apache-2.0·JSON
ImageMultimodalText

llava-en-zh-300k

This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…

100K–1M·Apache-2.0·Parquet
ImageMultimodalTabular

MVBench

MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…

1K–10K·MIT·JSON
ImageMultimodalText

pixmo-docs

PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…

100K–1M·ODC-BY·Parquet
ImageMultimodalText

PhyX

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…

10K–100K·MIT·Parquet
Text

OpenDocVQA

Dataset Card for OpenDocVQA This is a training and evaluation QA data file for VDocRAG, a new RAG…

10K–100K·Parquet
ImageMultimodalText

VLAA-Thinking

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…

<1K·Apache-2.0·Images (folder)
MultimodalTextVideo

VideoMarathon

Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…

1M–10M·Apache-2.0·Parquet
Text

synthvision-validated-qwen-by-kimi

synthvision-validated-qwen-by-kimi Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate) Records: 55,359 About Cross-validated subset from…

100K–1M·Apache-2.0·Parquet
MultimodalTextVideo

Video-opd-Dataset

Video-opd-Dataset A video temporal grounding dataset with 2,500 samples sourced from TimeLens-100K. Dataset Description This dataset contains video…

1K–10K·Apache-2.0·Parquet
Advertisement