Skip to content
Advertisement
Text

Turkish Corpus 100B

Turkish Corpus 100B

Turkish Corpus 100B (TC-100B) Dataset Summary The Turkish Corpus 100B (TC-100B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Turkish. Comprising approximately 105 Billion tokens (measured with Qwen/Llama3 tokenizer), it represents one of the largest open resources for Turkish LLM pretraining. The dataset is engineered for a two-stage training pipeline: Pretrain Subset (~103B Tokens): A diverse mix of… See the full description on the dataset page:

Source: Hugging Face Hub (hasankursun/turkish-corpus-100b). Metadata imported from the dataset’s Hub tags.

Advertisement