Skip to content
Advertisement
AudioMultimodalText

Omnilingual ASR Corpus

Omnilingual ASR Corpus

Meta Omnilingual ASR Corpus The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models. Data schema { language : “lij Latn”, iso 639 3 : “lij”, iso 15924 : “Latn”, glottocode :… See the full description on the dataset page:

Source: Hugging Face Hub (facebook/omnilingual-asr-corpus). Metadata imported from the dataset’s Hub tags.

Advertisement