Core-VIIRS-Nighttime-Light
Major TOM Core VIIRS Nighttime Light Annual radiance composites from the Visible Infrared Imaging Radiometer Suite (VIIRS) Day/Night…
Dataset Card for OpusParaCrawl Dataset Summary Parallel corpora from Web Crawls collected in the ParaCrawl project. Tha dataset contains: 42 languages, 43 bitexts total number of files: 59,996 total number of tokens: 56.11G total number of sentence fragments: 3.13G To load a language pair which isn’t part of the config, all you need to do is specify the language code as pairs, e.g. dataset = load dataset(“opus paracrawl”, lang1=”en”, lang2=”so”) You can find the valid… See the full description on the dataset page: paracrawl.
Source: Hugging Face Hub (Helsinki-NLP/opus_paracrawl). Metadata imported from the dataset’s Hub tags.