Skip to content
Advertisement

1M–10M

329 results

AudioMultimodalText

Galgame-VisualNovel-Reupload

Galgame VisualNovel Reupload This repository is a reupload of the visual novel dataset OOPPEENN/56697375616C4E6F76656C5F44617461736574. The goal of…

1M–10M·Custom / Research-only·Parquet
AudioMultimodalText

legco-speech

香港立法會會議語音數據集 本數據集係由香港立法會會議製成嘅大規模語音數據集。原始錄音總時長 22,196 個鐘,切分語音後總時長 20,471 個鐘。數據集分兩個子集,raw同segmented,分別為原始錄音同VAD識別切分後嘅語音。 數據集製作流程 先去香港特別行政區立法會…

1M–10M·CC0·Parquet
AudioMultimodalText

echo-clones-4m-en

echo-clones-4m-en ~4 M English TTS clone utterances generated with EchoTTS (jordand/echo-tts-base). Sample rate: 44 100 Hz, 16-bit PCM…

1M–10M·Apache-2.0·Parquet
Audio

advanced-soundscapes-stage-1

Advanced Soundscapes Stage 1 — Raw Components (5M) This dataset contains Stage 1 output from the LAION Universal…

1M–10M·CC-BY
AudioMultimodalText

BigAudioDataset

AstraMindAI/BigAudioDataset Dataset Description AstraMindAI/BigAudioDataset is a large-scale, multilingual dataset designed for a wide range of audio…

1M–10M·Apache-2.0·Arrow
AudioMultimodalText

IndicVoices

IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates 23 December 2025 We now have…

1M–10M·CC-BY·Parquet
Advertisement