Skip to content
Advertisement

10K–100K

748 results

AudioMultimodalText

AISHELL-3

AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…

10K–100K·Apache-2.0
AudioMultimodalText

opendata-iisys-hui

HUI-Audio-Corpus-German Dataset Overview The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the…

10K–100K·MIT·Parquet
AudioMultimodalText

voicebench

License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in…

10K–100K·Apache-2.0·Parquet
AudioMultimodalText

UrbanSound8K

UrbanSound8K This is an audio classification dataset for Sound Event Classification. Classes = 10 , Split = Ten-Fold…

10K–100K·MIT·CSV
AudioMultimodalText

floras

FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS…

10K–100K·CC-BY·Parquet
Audio

central-kurdish-audiobook-raw

Central Kurdish Audiobook Raw Audio Collection Overview This repository contains a large collection of raw Central Kurdish (Sorani…

10K–100K·Audio (folder)
Audio

FSD50k

Freesound Dataset 50k (FSD50K) Important This data set is a copy from the original one located at Zenodo.…

10K–100K·CC-BY
AudioMultimodalText

synthetic-asr-hi

Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
AudioMultimodalVideo

JavisInst-Omni

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…

10K–100K·Apache-2.0
AudioMultimodalText

synthetic-asr-zh

Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…

10K–100K·JSON
Advertisement