Skip to content
Advertisement
Text

ParaCrawl_Context

ParaCrawl Context

Dataset Card for ParaCrawl Context This is a dataset for document-level machine translation introduced in the ACL 2024 paper Document-Level Machine Translation with Large-Scale Public Parallel Data. It is a dataset consisting of parallel sentence pairs from the ParaCrawl dataset along with corresponding preceding context extracted from the webpages the sentences were crawled from. Dataset Details Dataset Description This dataset adds document-level… See the full description on the dataset page: context.

Source: Hugging Face Hub (Proyag/paracrawl_context). Metadata imported from the dataset’s Hub tags.

Advertisement