Text
2,175 results
smollm-corpus
SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…
webdev-arena-preference-10k
WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details…
fineweb-edu-translated
Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse Vchitect Team1 1Shanghai Artificial Intelligence Laboratory Paper Project Page Data Overview The Vchitect-T2V-Dataverse is the…
Crowd Recital Yiddish – Source Dataset
Crowd Recital Yiddish - Source Dataset
LongBench-v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…
LLaVA-Video-178K
Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only…
yodas2_sidon
YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis…