Skip to content
Advertisement
Text

nmt-parallel-corpus

Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned platform for…

Neural Machine Translation parallel corpora Introduction We use OpusTools to extract resources from the OPUS project, a renowned platform for parallel corpora, and create a multilingual dataset. Specifically, we collect the parallel corpora from prominent projects within OPUS, including NLLB, CCMatrix, and OpenSubtitles. This comprehensive data collection process results in a corpus of more than 3T, covering 60 languages and over 1900 language pairs.… See the full description on the dataset page:

Source: Hugging Face Hub (liboaccn/nmt-parallel-corpus). Metadata imported from the dataset’s Hub tags.

Advertisement