Multilingual
575 results
Parallel Corpus from United Nations Digital Library Official Document System aligned by file
Parallel Corpus from United Nations Digital Library Official Document System aligned by file
Banque_Sonore_Dialectes_Bretons
!NOTE Dataset origin: Description Issue du site Banque Sonore des Dialectes Bretons Présentation du projet La Banque Sonore…
DS-TF2-RO-3M — 3M English→Romanian Fable Translations
DS-TF2-RO-3M — 3M English→Romanian Fable Translations
bhasha-sft
Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual…
LightNovel5000
Light novels translated in Chinese - crawled from public websites that do not prohibit crawlers 脚盆轻小说汉化 - 从未禁止爬虫的公共网站爬取…
The Vuk’uzenzele South African Multilingual Corpus
The Vuk'uzenzele South African Multilingual Corpus
IndonesianNMT
This dataset is used on the paper "Replicable Benchmarking of Neural Machine Translation (NMT) on Low-Resource Local Languages…
tanglish-tamil
Translation of Tanglish to tamil Source: karky.in To use python import datasets s = datasets.load dataset('Deepakvictor/tanglish-tamil') print(s)…
skr_trans_distill
Dataset Card for skr trans distill 本数据集由 SakuraLLM 大模型生成机器翻译结果,主要用于模型蒸馏训练。数据集包含日文原文及其对应的中文翻译,适用于日译中任务的模型训练与蒸馏。 Dataset Details Dataset Description…
Synthetic Parallel EN-LG (external)
Synthetic Parallel EN-LG (external)
LAION-COCO translated to 200 languages
LAION-COCO translated to 200 languages
TED Translation Decision Dataset
TED Translation Decision Dataset
WMT20 – MultiLingual Quality Estimation (MLQE) Task2
WMT20 - MultiLingual Quality Estimation (MLQE) Task2
SETimes – A Parallel Corpus of English and South-East European Languages
SETimes – A Parallel Corpus of English and South-East European Languages
COCO-35L
Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…