Datasets
1,940 results
reazon-speech-v2-denoised-same-speaker-pair
Reazon Speech v2 DENOISED Same Speaker Pair prj-beatrice/reazon-speech-v2-denoised-titanet-embeddings のコサイン類似度の総和がなるべく大きくなるように…
AppTek Call-Center Dialogues
AppTek Call-Center Dialogues
Tarteel AI – EveryAyah Dataset
Tarteel AI - EveryAyah Dataset
AudioMarathon
AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings.…
audioset-16khz-wds
AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527…
ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
avspeech-visual-audio
AVSpeech Video + Audio A restructured subset of the AVSpeech dataset with separated media streams and derived identifiers.…
meow-10k
Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…
Vietnamese Restaurant Order Speech
Vietnamese Restaurant Order Speech
libritts_r
Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…
Bengali Telecom Customer Care Synthetic Speech v2
Bengali Telecom Customer Care Synthetic Speech v2
MMAE
MMAE: A Massive Multitask Audio Editing Benchmark 📖 arXiv 🎬 MMAE Demo Video 🛠️ GitHub Code 🔊 HuggingFace…
AISHELL-3
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…
cml-tts-filtered
Dataset Card for Filtred and CML-TTS This dataset is a filtred version of a CML-TTS 1 . CML-TTS…