Skip to content
Advertisement

Datasets

3,216 results

AudioMultimodalText

MMAE

MMAE: A Massive Multitask Audio Editing Benchmark 📖 arXiv 🎬 MMAE Demo Video 🛠️ GitHub Code 🔊 HuggingFace…

1K–10K·Audio (folder)
AudioMultimodalText

AISHELL-3

AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…

10K–100K·Apache-2.0
AudioMultimodalText

MRSAudio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…

100K–1M·CC-BY·CSV
AudioMultimodalText

synthetic_vocal_bursts

This repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. We captioned…

100K–1M·Apache-2.0·WebDataset
AudioMultimodalText

Galgame-VisualNovel-Reupload

Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…

1M–10M·Custom / Research-only·Parquet
Audio

LibriBrain

LibriBrain (Sherlock Holmes 1–7) Paper Code This repository contains the LibriBrain data organised by book: MEG recordings (.h5),…

CC-BY-NC
AudioMultimodalText

legco-speech

香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…

1M–10M·CC0·Parquet
AudioMultimodalVideo

rh20t_cfg1

rh20t cfg1 — RH20T → LeRobot v3 (flexiv) Unofficial LeRobot Dataset v3 reformatting of an RH20T config. RGB…

1K–10K·Custom / Research-only·Audio (folder)
AudioMultimodalText

opendata-iisys-hui

HUI-Audio-Corpus-German Dataset Overview The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the…

10K–100K·MIT·Parquet
Advertisement