Multitask-National-Speech-Corpus-v1
Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and…
2,368 results
Multitask-National-Speech-Corpus (MNSC v1) is derived from IMDA's NSC Corpus. MNSC is a multitask speech understanding dataset derived and…
Tahoe-100M Tahoe-100M is a giga-scale single-cell perturbation atlas consisting of over 100 million transcriptomic profiles from 50 cancer…
CodeParrot 🦜 Dataset Cleaned What is it? A dataset of Python files from Github. This is the deduplicated…
EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…
This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included…
SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…
Crowd Recital Yiddish - Source Dataset
Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only…
YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis…
ICLR 2026 DCASE 2026 Training Set AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset,…
Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…
InsViE-1M: Effective Instruction-based Video Editing with Elaborate Dataset Construction Citation If you find this work helpful, please consider…