Skip to content
Advertisement
Text

Bulgarian Corpus 33B

Bulgarian Corpus 33B

Bulgarian Corpus 33B (BC-33B) Dataset Summary The Bulgarian Corpus 33B (BC-33B) is a massive-scale, deduplicated, and cleaned dataset designed for training Foundation Models in Bulgarian. Comprising approximately 33.4 Billion tokens (measured with Qwen 2.5/Llama-3 tokenizer), it represents one of the largest open-source resources for Bulgarian LLM pretraining. The dataset is engineered for a modern two-stage training pipeline: Pretrain Subset (~29.3B Tokens): A… See the full description on the dataset page:

Source: Hugging Face Hub (hasankursun/bulgarian-corpus-33b). Metadata imported from the dataset’s Hub tags.

Advertisement