databricks-dolly-15k
Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several…
3,216 results
Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several…
EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…
Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the…
This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included…
SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…
WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details…
Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…
Vchitect-T2V-Dataverse Vchitect Team1 1Shanghai Artificial Intelligence Laboratory Paper Project Page Data Overview The Vchitect-T2V-Dataverse is the…
Crowd Recital Yiddish - Source Dataset