muaalem-annotated-v3
قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء…
2,175 results
قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء…
Speech Recognition Alignment Dataset
Jalak Indonesian Multi-Speaker TTS
License The dataset is available under the Apache 2.0 license. Citation If you use the VoiceBench dataset in…
UrbanSound8K This is an audio classification dataset for Sound Event Classification. Classes = 10 , Split = Ten-Fold…
MLS-Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for…
Bo du lieu giong noi tieng Viet
Daily-Omni (QA repackaged for lmms-eval)
FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS…
(card and dataset copied from This dataset contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10…
CineAudioSynth Synthetic cinematic audio for source separation. 453 scenes, ~22.8 h, 48 kHz / 16-bit / stereo WAV.…
Dataset Card for μ-Bench (Leaderboard Code) μ-Bench is a multilingual transcription benchmark built from real customer-service phone conversations.…
Multilingual MFA-Aligned Speech Dataset
Police Scanner Audio Dataset A comprehensive collection of police and emergency services radio communications from multiple US cities,…
Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned…