Core-VIIRS-Nighttime-Light
Major TOM Core VIIRS Nighttime Light Annual radiance composites from the Visible Infrared Imaging Radiometer Suite (VIIRS) Day/Night…
Dutch Corpus 200B (DC-200B) Dataset Summary The Dutch Corpus 200B (DC-200B) is the largest open-source, deduplicated, and professionally cleaned dataset designed for training Foundation Models in the Dutch language. Comprising approximately 202 Billion tokens (measured with Qwen 2.5 tokenizer), it bridges the gap between high-resource English models and the Dutch ecosystem. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~195B… See the full description on the dataset page:
Source: Hugging Face Hub (hasankursun/dutch-corpus-200b). Metadata imported from the dataset’s Hub tags.