Skip to content
Advertisement
Text

Africa Corpus

Africa Corpus

Africa Corpus Verse-aligned text for 693 African languages, plus several world languages, for building monolingual and parallel corpora. Every language is aligned on a shared verse key, so any single language can be pulled on its own or any two joined into a parallel corpus: Monolingual corpus for any single language African ↔ English (English is the default pair) African ↔ African (e.g. Twi ↔ Yoruba, Hausa ↔ Amharic) African ↔ other language (French, Arabic, Chinese… See the full description on the dataset page:

Source: Hugging Face Hub (AfriSpeech/africa-corpus). Metadata imported from the dataset’s Hub tags.

Advertisement