audioset-16khz-wds
AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527…
2,175 results
AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527…
ASMR-Archive-Processed-SFW
AVSpeech Video + Audio A restructured subset of the AVSpeech dataset with separated media streams and derived identifiers.…
Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…
Vietnamese Restaurant Order Speech
Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…
Bengali Telecom Customer Care Synthetic Speech v2
MMAE: A Massive Multitask Audio Editing Benchmark 📖 arXiv 🎬 MMAE Demo Video 🛠️ GitHub Code 🔊 HuggingFace…
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…
Dataset Card for Filtred and CML-TTS This dataset is a filtred version of a CML-TTS 1 . CML-TTS…
Dataset Card for "speech robust bench" More Information needed
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…
This repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. We captioned…
shkolkovo-bobr.video-webinars-audio Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars…
FreeSound.org LAION-640k Dataset (16 KHz)
Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…
香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…
Vāgdhenu — Sanskrit Chant Corpus
HUI-Audio-Corpus-German Dataset Overview The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the…
common-voice-asr-clean Filtered ASR dataset. Samples with <3 words, repetitive tokens, or chat token leaks removed.