quickmt-train.hu-en
quickmt hu-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…
2,594 results
quickmt hu-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…
ScreenTalk JA2ZH-XS ScreenTalk JA2ZH-XS is a paired dataset of Japanese speech and Chinese translated text released by DataLabX.…
English-ASL Gloss Parallel Corpus 2012
English–Zomi Parallel Corpus (1.78M)
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…
!NOTE Dataset origin: Context TED is devoted to spreading powerful ideas in just about any topic. These datasets…
ChemrXiv Pdf Introducing ChemrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024.…
Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…
Dataset Card for Planetarium🪐 Planetarium🪐 is a dataset and benchmark for assessing LLMs in translating natural language descriptions…
English-Polish MetricX-filtered Parallel Sentences with Qwen3 Embeddings
联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score…