Skip to content
Advertisement

Apache-2.0

807 results

Image

minimind-v_dataset

Ⅰ 数据集 本轮训练用到的图文数据全部来自 ALLaVA-4V 系列。 相比以往从几份 LLaVA 衍生集拼接得到的数据,ALLaVA-4V 的质量更整齐、中英双语原生对照,细粒度描述也更充分。 它由两个子源构成:一份是 LAION 里挑出来的高质量图片(自然图像为主),一份是 VFLAN…

<1K·Apache-2.0·Images (folder)
Image

VisuLogic

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models A Challenging Visual-centric Benchmark for Evaluating…

1K–10K·Apache-2.0·Images (folder)
ImageMultimodalText

GenEvolve-Data-Bench

GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose…

10K–100K·Apache-2.0·Parquet
MultimodalTextVideo

VideoKR-Eval

VideoKR-Eval 📄 ArXiv  |  💻 Code  |  🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…

1K–10K·Apache-2.0·JSON
ImageMultimodalText

llava-en-zh-300k

This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…

100K–1M·Apache-2.0·Parquet
ImageMultimodalText

VLAA-Thinking

SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…

<1K·Apache-2.0·Images (folder)
MultimodalTextVideo

VideoMarathon

Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…

1M–10M·Apache-2.0·Parquet
Text

synthvision-validated-qwen-by-kimi

synthvision-validated-qwen-by-kimi Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate) Records: 55,359 About Cross-validated subset from…

100K–1M·Apache-2.0·Parquet
MultimodalTextVideo

Video-opd-Dataset

Video-opd-Dataset A video temporal grounding dataset with 2,500 samples sourced from TimeLens-100K. Dataset Description This dataset contains video…

1K–10K·Apache-2.0·Parquet
ImageMultimodalText

ChartVerse-SFT-600K

ChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…

100K–1M·Apache-2.0·Parquet
MultimodalTextVideo

videomarathon

Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…

1M–10M·Apache-2.0·Parquet
MultimodalTextVideo

ReVSI

ICML 2026 Yiming Zhang1 , Jiacheng Chen1 , Jiaqi Tan1, Yongsen Mao2, Wenhu Chen3, Angel X. Chang1,4 1…

10K–100K·Apache-2.0·Parquet
Advertisement