Skip to content
Advertisement
Text

Alexandria Multudialectal Arabic Conversational Dataset for Machine Translation

Alexandria Multudialectal Arabic Conversational Dataset for Machine Translation

Dataset Card for Alexandria Alexandria covers 13 Arab countries, 11 domains, and 107K community-driven samples. Alexandria is a multi-domain English↔Dialectal Arabic machine translation dataset designed for culturally inclusive, dialect-aware NLP and LLM evaluation. It pairs English multi-turn conversations with human-translated dialectal Arabic from 13 Arab countries, enriched with sub-dialect metadata (based on city-level information), domain labels, persona roles… See the full description on the dataset page:

Source: Hugging Face Hub (UBC-NLP/alexandria). Metadata imported from the dataset’s Hub tags.

Advertisement