Skip to content
Advertisement

SmolTalk Dataset description This is a synthetic dataset designed for supervised finetuning (SFT) of LLMs. It was used to build SmolLM2-Instruct family of models and contains 1M samples. More details in our paper During the development of SmolLM2, we observed that models finetuned on public SFT datasets underperformed compared to other models with proprietary instruction datasets. To address this gap, we created new synthetic datasets… See the full description on the dataset page:

Source: Hugging Face Hub (HuggingFaceTB/smoltalk). Metadata imported from the dataset’s Hub tags.

Advertisement