Skip to content
Advertisement
Text

DocHPLT

DocHPLT

DocHPLT: A Massively Multilingual Document-Level Translation Dataset Existing document-level machine translation resources are only available for a handful of languages, mostly high-resourced ones. To facilitate the training and evaluation of document-level translation and, more broadly, long-context modeling for global communities, we create DocHPLT, the largest publicly available document-level translation dataset to date. It contains 124 million aligned document pairs across 50… See the full description on the dataset page:

Source: Hugging Face Hub (HPLT/DocHPLT). Metadata imported from the dataset’s Hub tags.

Advertisement