meow-10k
Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…
807 results
Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…
This repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. We captioned…
Speech Recognition Alignment Dataset
License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in…
Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned…
echo-clones-4m-en ~4 M English TTS clone utterances generated with EchoTTS (jordand/echo-tts-base). Sample rate: 44 100 Hz, 16-bit PCM…
Irodori TTS Clones v2 (3.29M) 3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2, using the 10,000 reference voices from…
AISHELL-4 Identifier: SLR111 Summary: A Free Mandarin Multi-channel Meeting Speech Corpus, provided by Beijing Shell Shell Technology Co.,Ltd…
BEAT2-English + additional annotations (RAG-Gesture / MIBURI)
French Conversational TTS Dataset Dataset Description This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's…
Reachy Mini Emotions Library
TTS Pretrain Clones (3M) — with DNSMOS This is SynDataLab/tts-pretrain-clones-3m with an added per-utterance dnsmos column (DNSMOS P.835…
AstraMindAI/BigAudioDataset Dataset Description AstraMindAI/BigAudioDataset is a large-scale, multilingual dataset designed for a wide range of audio…
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…
Marco-LongSpeech Dataset Marco-LongSpeech is a multi-task long speech understanding dataset containing 8 different speech understanding tasks…
Chinese Fineweb Edu Dataset V2.1 中文 English OpenCSG Community 👾github wechat Twitter 📖Technical Report The Chinese Fineweb Edu…
LOTSA Data The Large-scale Open Time Series Archive (LOTSA) is a collection of open time series datasets for…