Skip to content
Advertisement
Text

TaPaCo Corpus

TaPaCo Corpus

Dataset Card for TaPaCo Corpus Dataset Summary A freely available paraphrase corpus for 73 languages extracted from the Tatoeba database. Tatoeba is a crowdsourcing project mainly geared towards language learners. Its aim is to provide example sentences and translations for particular linguistic constructions and words. The paraphrase corpus is created by populating a graph with Tatoeba sentences and equivalence links between sentences “meaning the same thing”. This… See the full description on the dataset page:

Source: Hugging Face Hub (community-datasets/tapaco). Metadata imported from the dataset’s Hub tags.

Advertisement