Skip to content
Advertisement
Text

Vast Urdu Parallel Corpus

Vast Urdu Parallel Corpus

Vast Urdu Parallel Corpus Dataset Description Vast-Urdu is a large-scale collection of parallel text corpora specifically filtered to support Urdu (UR) language research. This dataset was extracted from the liboaccn/nmt-parallel-corpus to provide a dedicated resource for Neural Machine Translation (NMT), cross-lingual understanding, and token-classification tasks involving Urdu. Source Data The data is sourced from a massive web-scale crawl, containing… See the full description on the dataset page:

Source: Hugging Face Hub (Humair332/Vast-Urdu). Metadata imported from the dataset’s Hub tags.

Advertisement