Skip to content
Advertisement

10K–100K

748 results

Text

AfriDocMT

data ├── document │ ├── Health │ │ ├── dev.csv │ │ ├── test.csv │ │ └── train.csv…

10K–100K·CSV
Text

OpusGnome

OpusGnome

10K–100K·Custom / Research-only·Parquet
Text

NL2SH-ALFA

Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…

10K–100K·MIT·CSV
ImageMultimodalTabular

ChEBI-20-MM

ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is…

10K–100K·MIT·CSV
ImageMultimodalText

Flickr30k

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

10K–100K·CC-BY·Parquet
Text

ECDC

ECDC

10K–100K·Custom / Research-only
AudioMultimodalText

MultiMed-ST

MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation 📘 EMNLP 2025 Khai Le-Duc , Tuyen Tran , Bach Phan…

10K–100K·MIT·Parquet
Text

IndonesianNMT

This dataset is used on the paper "Replicable Benchmarking of Neural Machine Translation (NMT) on Low-Resource Local Languages…

10K–100K
Text

TinyStories-Multilingual

Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short,…

10K–100K·Apache-2.0·JSON
Text

wmt24pp

WMT24++ This repository contains the human translation and post-edit data for the 55 en- xx language pairs released…

10K–100K·Apache-2.0·JSON
Text

PHINC

Abstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities,…

10K–100K·CC-BY·CSV
Text

multi30k

Multi30k This dataset contains the "multi30k" dataset, which is the "task 1" dataset from here. Each example consists…

10K–100K·JSON
ImageMultimodalText

BEAF

BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…

10K–100K·CC-BY·Parquet
ImageMultimodalText

GQA-ru

GQA-ru This is a translated version of original GQA dataset and stored in format supported for lmms-eval pipeline.…

10K–100K·Apache-2.0·Parquet
Advertisement