Persian Common Voice Clean Dataset This dataset is a cleaned and prepared subset of the Persian (فارسی – fa) portion of Mozilla Common Voice Scripted Speech, based on cv-corpus-26.0-2026-06-12. The cleaned release contains 34,134 audio clips, representing approximately 43.105 hours of speech, equal to 2,586.303 minutes. The clips are associated with approximately 34,134 validated Persian sentences and come from 3,791 speakers. The original Persian Common Voice release contains… See the full description on the dataset page:
Source: Hugging Face Hub (pymmdrza/Common-Voice-Speech-26.0-Persian-Clean). Metadata imported from the dataset’s Hub tags.