Skip to content
Advertisement
Text

OPUS EUconst

OPUS EUconst

Dataset Card for OPUS EUconst Dataset Summary A parallel corpus collected from the European Constitution. EUconst’s Numbers: Languages: 21 Bitexts: 210 Number of files: 986 Number of tokens: 3.01M Sentence fragments: 0.22M Supported Tasks and Leaderboards The underlying task is machine translation. Languages The languages in the dataset are: Czech (cs) Danish (da) German (de) Greek (el) English (en) Spanish (es) Estonian (et) Finnish (fi) French… See the full description on the dataset page:

Source: Hugging Face Hub (Helsinki-NLP/euconst). Metadata imported from the dataset’s Hub tags.

Advertisement