Skip to content
Advertisement

Text

2,175 results

AudioMultimodalText

audioset-16khz-wds

AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527…

1M–10M·WebDataset
AudioMultimodalText

meow-10k

Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…

10K–100K·Apache-2.0·JSON
AudioMultimodalText

libritts_r

Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…

100K–1M·CC-BY·Parquet
AudioMultimodalText

MMAE

MMAE: A Massive Multitask Audio Editing Benchmark 📖 arXiv 🎬 MMAE Demo Video 🛠️ GitHub Code 🔊 HuggingFace…

1K–10K·Audio (folder)
AudioMultimodalText

AISHELL-3

AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…

10K–100K·Apache-2.0
AudioMultimodalText

MRSAudio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…

100K–1M·CC-BY·CSV
AudioMultimodalText

synthetic_vocal_bursts

This repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. We captioned…

100K–1M·Apache-2.0·WebDataset
AudioMultimodalText

Galgame-VisualNovel-Reupload

Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…

1M–10M·Custom / Research-only·Parquet
AudioMultimodalText

legco-speech

香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…

1M–10M·CC0·Parquet
AudioMultimodalText

opendata-iisys-hui

HUI-Audio-Corpus-German Dataset Overview The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the…

10K–100K·MIT·Parquet
Advertisement