Skip to content
Advertisement

Translation

253 results

Text

OpusGnome

OpusGnome

10K–100K·Custom / Research-only·Parquet
Text

Parallel_Dataset

Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:

1M–10M·MIT·JSON
Text

NL2SH-ALFA

Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…

10K–100K·MIT·CSV
ImageMultimodalTabular

ChEBI-20-MM

ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is…

10K–100K·MIT·CSV
MultimodalTabularText

JMedBench

Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any…

100K–1M·JSON
ImageMultimodalText

Flickr30k

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

10K–100K·CC-BY·Parquet
Text

UnGa

UnGa

1M–10M·Custom / Research-only·Parquet
Text

NTREX

Dataset Description NTREX -- News Test References for MT Evaluation from English into a total of 128 target…

100K–1M·CC-BY-SA·Text (raw)
Text

OPUS-100

OPUS-100

10M–100M·Custom / Research-only·Parquet
Text

WenYanWen_English_Parallel

Dataset Card for WenYanWen English Parallel Dataset Summary The WenYanWen English Parallel dataset is a multilingual parallel corpus…

1M–10M·MIT·Parquet
Text

mmarco-contrastive

mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to…

100K–1M·Apache-2.0·Parquet
Text

ECDC

ECDC

10K–100K·Custom / Research-only
AudioMultimodalText

MultiMed-ST

MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation 📘 EMNLP 2025 Khai Le-Duc , Tuyen Tran , Bach Phan…

10K–100K·MIT·Parquet
Advertisement