Skip to content
Advertisement
Text

Shaistagi (شائستگی) Clean Urdu Mega-Dataset

Shaistagi (شائستگی) Clean Urdu Mega-Dataset

Shaistagi (شائستگی) Clean Urdu Mega-Dataset The largest and most comprehensive cleaned Urdu NLP dataset collection 📊 Dataset Overview Shaistagi Clean is one of the most comprehensive, multi-task Urdu NLP collections available. It aggregates high-quality, cleaned data for pre-training, instruction-tuning, and specialized downstream tasks. Key Statistics Metric Value Total Rows ~16.2 Million Total Tokens ~1.22 Billion Total Characters… See the full description on the dataset page: clean.

Source: Hugging Face Hub (ReySajju742/shaistagi_clean). Metadata imported from the dataset’s Hub tags.

Advertisement