AISHELL-3
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…
748 results
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…
HUI-Audio-Corpus-German Dataset Overview The HUI-Audio-Corpus-German is a high-quality Text-To-Speech (TTS) dataset developed by researchers at the…
Jalak Indonesian Multi-Speaker TTS
License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in…
UrbanSound8K This is an audio classification dataset for Sound Event Classification. Classes = 10 , Split = Ten-Fold…
FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS…
NISR — Neural Inverse Sound Rendering Dataset
Central Kurdish Audiobook Raw Audio Collection Overview This repository contains a large collection of raw Central Kurdish (Sorani…
A Multilingual Speech Dataset for SLU and Beyond
English-Centric Multilingual Audio Dataset
Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation HomePage Paper GitHub TL;DR We introduce JavisGPT, a…
AVQA JSONL (Audio Multiple-Choice QA)
Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…