Skip to content
Advertisement

Audio

514 results

AudioMultimodalText

meow-10k

Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…

10K–100K·Apache-2.0·JSON
AudioMultimodalText

libritts_r

Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…

100K–1M·CC-BY·Parquet
AudioMultimodalText

MMAE

MMAE: A Massive Multitask Audio Editing Benchmark 📖 arXiv 🎬 MMAE Demo Video 🛠️ GitHub Code 🔊 HuggingFace…

1K–10K·Audio (folder)
AudioMultimodalText

AISHELL-3

AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…

10K–100K·Apache-2.0
AudioMultimodalText

MRSAudio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…

100K–1M·CC-BY·CSV
AudioMultimodalText

synthetic_vocal_bursts

This repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. We captioned…

100K–1M·Apache-2.0·WebDataset
AudioMultimodalText

Galgame-VisualNovel-Reupload

Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…

1M–10M·Custom / Research-only·Parquet
Audio

LibriBrain

LibriBrain (Sherlock Holmes 1–7) Paper Code This repository contains the LibriBrain data organised by book: MEG recordings (.h5),…

CC-BY-NC
AudioMultimodalText

legco-speech

香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…

1M–10M·CC0·Parquet
AudioMultimodalVideo

rh20t_cfg1

rh20t cfg1 — RH20T → LeRobot v3 (flexiv) Unofficial LeRobot Dataset v3 reformatting of an RH20T config. RGB…

1K–10K·Custom / Research-only·Audio (folder)
Advertisement