ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
329 results
ASMR-Archive-Processed-SFW
AVSpeech Video + Audio A restructured subset of the AVSpeech dataset with separated media streams and derived identifiers.…
Dataset Card for Filtred and CML-TTS This dataset is a filtred version of a CML-TTS 1 . CML-TTS…
Dataset Card for "speech robust bench" More Information needed
Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…
香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…
common-voice-asr-clean Filtered ASR dataset. Samples with <3 words, repetitive tokens, or chat token leaks removed.
Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned…
echo-clones-4m-en ~4 M English TTS clone utterances generated with EchoTTS (jordand/echo-tts-base). Sample rate: 44 100 Hz, 16-bit PCM…
Irodori TTS Clones v2 (3.29M) 3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2, using the 10,000 reference voices from…
Advanced Soundscapes Stage 1 — Raw Components (5M) This dataset contains Stage 1 output from the LAION Universal…
TTS Pretrain Clones (3M) — with DNSMOS This is SynDataLab/tts-pretrain-clones-3m with an added per-utterance dnsmos column (DNSMOS P.835…
LIEPA-3 Lithuanian Speech Corpus
AstraMindAI/BigAudioDataset Dataset Description AstraMindAI/BigAudioDataset is a large-scale, multilingual dataset designed for a wide range of audio…
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates 23 December 2025 We now have…