Skip to content
Advertisement
MultimodalTabularText

FineWeb Tokenized (AnisoleAI)

FineWeb Tokenized (AnisoleAI)

FineWeb Tokenized 4 trillion tokens of the pre-tokenized data the 🌐 web has to offer What is it? This is a pre-tokenized version of the HuggingFaceFW/fineweb dataset (currently in-progress, tokenization of the ~15 trillion tokens corpus is ongoing). The data is being pre-processed and tokenized using the AnisoleAI BPE tokenizer (52,022 vocabulary size) and packed into compact uint16 Parquet shards. By distributing the pre-tokenized corpus, we eliminate… See the full description on the dataset page:

Source: Hugging Face Hub (anisoleai/fineweb-tokenized). Metadata imported from the dataset’s Hub tags.

Advertisement