Skip to content
Advertisement
Text

OPUS

Collection of OPUS Corpus from has been collected. The following corpora have been included: UNPC GlobalVoices TED2020 News-Commentary WikiMatrix…

Collection of OPUS Corpus from has been collected. The following corpora have been included: UNPC GlobalVoices TED2020 News-Commentary WikiMatrix Tatoeba Europarl OpenSubtitles 25,000 samples (randomly sampled within the first 100,000 samples) per language pair of each corpus were collected, with no modification of data. Licenses OPUS @inproceedings{tiedemann2012parallel, title={Parallel data, tools and interfaces in OPUS.}… See the full description on the dataset page:

Source: Hugging Face Hub (wecover/OPUS). Metadata imported from the dataset’s Hub tags.

Advertisement