Text
2,175 results
OpenR1-Math-220k
OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with…
Multitask-National-Speech-Corpus-v1
Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and…
Tahoe-100M
Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer…
MetaMathQA
View the project page: see our paper at Note All MetaMathQA data are augmented from the training sets…
codeparrot-clean
CodeParrot 🦜 Dataset Cleaned What is it? A dataset of Python files from Github. This is the deduplicated…
databricks-dolly-15k
Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several…
EuroSpeech
EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…
nemotron-cc-translated
Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the…
covost2
This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included…