Multilingual
575 results
Text
OPUS Europarl (European Parliament Proceedings Parallel Corpus)
OPUS Europarl (European Parliament Proceedings Parallel Corpus)
100M–1B·Custom / Research-only·Parquet
Image
minimind-v_dataset
Ⅰ 数据集 本轮训练用到的图文数据全部来自 ALLaVA-4V 系列。 相比以往从几份 LLaVA 衍生集拼接得到的数据,ALLaVA-4V 的质量更整齐、中英双语原生对照,细粒度描述也更充分。 它由两个子源构成:一份是 LAION 里挑出来的高质量图片(自然图像为主),一份是 VFLAN…
<1K·Apache-2.0·Images (folder)
ImageMultimodalText
BoundingDocs
BoundingDocs 🔍 The largest spatially-annotated dataset for Document Question Answering Dataset Description BoundingDocs is a unified dataset for…
10K–100K·CC-BY·Parquet
ImageMultimodalTabular
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench (Public Ladder)
1K–10K·CC-BY-NC·JSON
ImageMultimodalText
DCVLM-Baseline (200B tokens)
DCVLM-Baseline (200B tokens)
10K–100K·Custom / Research-only·Parquet
ImageMultimodalText
llava-en-zh-300k
This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…
100K–1M·Apache-2.0·Parquet
MultimodalTextVideo
optimism_train
Optimism Bias World Model Benchmark — Training Data Training data for VLM-as-judge models evaluating world model predictions. Structure…
<1K·MIT