Skip to content
Advertisement

1M–10M

329 results

Text

UltraLink

multi-lingual, knowledge-grounded, multi-round dialogue dataset and model Summary • Construction Process • Paper • UltraLink-LM • Github Dataset…

1M–10M·MIT·JSON
Text

OPUS DGT

OPUS DGT

1M–10M·Custom / Research-only·Parquet
Text

rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…

1M–10M·Apache-2.0·CSV
Text

Parallel_Dataset

Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:

1M–10M·MIT·JSON
Text

UnGa

UnGa

1M–10M·Custom / Research-only·Parquet
Text

WenYanWen_English_Parallel

Dataset Card for WenYanWen English Parallel Dataset Summary The WenYanWen English Parallel dataset is a multilingual parallel corpus…

1M–10M·MIT·Parquet
Advertisement