Skip to content
Advertisement
MultimodalTabularText

pile-deduped-pythia-preshuffled

This dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia (deduplicated) models. You can…

This dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia (deduplicated) models. You can find these models under the EleutherAI organisation, and they are also listed in my Memorisation-Profiles collection. This data is the same as the one found in EleutherAI/pile-deduped-pythia-preshuffled, but it is presented in a more manageable format. Instead of using the Megatron format used by the GPT-NeoX library, I have stored the data in a… See the full description on the dataset page:

Source: Hugging Face Hub (pietrolesci/pile-deduped-pythia-preshuffled). Metadata imported from the dataset’s Hub tags.

Advertisement