Skip to content
Advertisement
Text

UltraChat 200k

UltraChat 200k

Dataset Card for UltraChat 200k Dataset Description This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model. The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic: Selection of a subset of data for faster supervised fine tuning. Truecasing of the dataset, as we observed around 5% of… See the full description on the dataset page: 200k.

Source: Hugging Face Hub (HuggingFaceH4/ultrachat_200k). Metadata imported from the dataset’s Hub tags.

Advertisement