Skip to content
Advertisement

Dataset Card for SADA (Saudi Audio Dataset for Arabic) Dataset Summary The SADA dataset (Saudi Audio Dataset for Arabic) is a large-scale Arabic speech corpus designed to support the development of high-quality artificial intelligence models for Arabic speech processing. It contains over 667 hours of transcribed Arabic audio recordings, primarily featuring various Saudi dialects, and was curated in a collaboration between the National Center for Artificial… See the full description on the dataset page:

Source: Hugging Face Hub (MohamedRashad/SADA22). Metadata imported from the dataset’s Hub tags.

Advertisement