Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
🧬 Carbon Pretraining Corpus Description 173M DNA & RNA sequences · 1.1 trillion nucleotides — the DNA pretraining mixture used to train Carbon, a genomic foundation model. This dataset is a collection of data sources intended for training genomic foundation models, such as Carbon. It contains DNA and RNA sequences spanning eukaryote and prokaryote species. Across the four main configs it totals 1.1 T DNA base pairs (180B tokens with Carbon’s 6-mer tokenizer). A… See the full description on the dataset page:
Source: Hugging Face Hub (HuggingFaceBio/carbon-pretraining-corpus). Metadata imported from the dataset’s Hub tags.