Skip to content
Advertisement

MediaSpeech MediaSpeech is a dataset of Arabic, French, Spanish, and Turkish media speech built with the purpose of testing Automated Speech Recognition (ASR) systems performance. The dataset contains 10 hours of speech for each language provided. The dataset consists of short speech segments automatically extracted from media videos available on YouTube and manually transcribed, with some pre-processing and post-processing. Baseline models and WAV version of the dataset can be… See the full description on the dataset page:

Source: Hugging Face Hub (ymoslem/MediaSpeech). Metadata imported from the dataset’s Hub tags.

Advertisement