Datasets
3,216 results
Russian Podcasts (unlabeled)
Russian Podcasts (unlabeled)
Danbooruwildcards
This is a set of wildcards for danbooru tags. Artist:Prompts for random artist styles, covering approximately 0.6M different…
Stanford Sentiment Treebank v2
Stanford Sentiment Treebank v2
Spotify Huge Track Analysis Dataset
Spotify Huge Track Analysis Dataset
Deep Dialogue (Orpheus TTS)
Deep Dialogue (Orpheus TTS)
OpenR1-Math-220k
OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with…
Multitask-National-Speech-Corpus-v1
Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and…
Tahoe-100M
Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer…
MetaMathQA
View the project page: see our paper at Note All MetaMathQA data are augmented from the training sets…
codeparrot-clean
CodeParrot 🦜 Dataset Cleaned What is it? A dataset of Python files from Github. This is the deduplicated…