Skip to content
Advertisement
Text

MultiUN (Multilingual Corpus from United Nation Documents)

MultiUN (Multilingual Corpus from United Nation Documents)

Dataset Card for OPUS MultiUN Dataset Summary The MultiUN parallel corpus is extracted from the United Nations Website , and then cleaned and converted to XML at Language Technology Lab in DFKI GmbH (LT-DFKI), Germany. The documents were published by UN from 2000 to 2009. This is a collection of translated documents from the United Nations originally compiled by Andreas Eisele and Yu Chen (see This corpus is available in all 6… See the full description on the dataset page:

Source: Hugging Face Hub (Helsinki-NLP/multiun). Metadata imported from the dataset’s Hub tags.

Advertisement