Skip to content
Advertisement
Text

Greek Corpus 150B

Greek Corpus 150B

Greek Corpus 150B A large-scale, deduplicated Greek (Modern Greek, el) text corpus for training and fine-tuning foundation models. It pairs a broad web/knowledge/formal-document pretrain layer with a multilingual-instruction SFT layer, all normalized to a single unified schema and globally deduplicated. This is part of an ongoing Global Corpus family of per-language foundation-model datasets (Dutch, Turkish, Bulgarian, Greek, …) built on a consistent architecture so that sources… See the full description on the dataset page:

Source: Hugging Face Hub (hasankursun/greek-corpus-150b). Metadata imported from the dataset’s Hub tags.

Advertisement