Skip to content
Advertisement
Text

OpusTedtalks

OpusTedtalks

Dataset Card for OpusTedtalks Dataset Summary This is a Croatian-English parallel corpus of transcribed and translated TED talks, originally extracted from The corpus is compiled by Željko Agić and is taken from provided under the CC-BY-NC-SA license. This corpus is sentence aligned for both language pairs. The documents were collected and aligned using the Hunalign algorithm. Supported Tasks and Leaderboards… See the full description on the dataset page: tedtalks.

Source: Hugging Face Hub (Helsinki-NLP/opus_tedtalks). Metadata imported from the dataset’s Hub tags.

Advertisement