Skip to content
Advertisement
MultimodalTabularText

GPT-NL_Public_Corpus

Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language…

Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code, and 48B German/Danish tokens. All data is sourced under permissive licensing and redistributed under a CC-BY license. For more details please refer to our Public Corpus article. Dataset… See the full description on the dataset page: Public Corpus.

Source: Hugging Face Hub (GPT-NL/GPT-NL_Public_Corpus). Metadata imported from the dataset’s Hub tags.

Advertisement