Skip to content
Advertisement
MultimodalTabularText

Carbon Pretraining Corpus

Carbon Pretraining Corpus

🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon’s 6-mer tokenizer). A… See the full description on the dataset page:

Source: Hugging Face Hub (HuggingFaceBio/carbon-pretraining-corpus). Metadata imported from the dataset’s Hub tags.

Advertisement