Skip to content
Advertisement

C4 Dataset Summary A colossal, cleaned version of Common Crawl’s web crawl corpus. Based on Common Crawl dataset: ” This is the processed version of Google’s C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page:

Source: Hugging Face Hub (allenai/c4). Metadata imported from the dataset’s Hub tags.

Advertisement