Skip to content
Advertisement
Text

Corpus-200B

WebOrganizer/Corpus-200B Paper Website GitHub This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with…

WebOrganizer/Corpus-200B Paper Website GitHub This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned with (1) RefinedWeb filters and (2) BFF deduplication. We provide the resulting 200B token corpus annotated with two quality scores, WebOrganizer domains, and k-means scores. Download the dataset by cloning the repository with Git LFS instead of HuggingFace’s load dataset(). The dataset has the following folder structure:… See the full description on the dataset page:

Source: Hugging Face Hub (WebOrganizer/Corpus-200B). Metadata imported from the dataset’s Hub tags.

Advertisement