Skip to content
Advertisement
MultimodalTabularText

Classifiers-Data

Text Quality Classifier Training Dataset This dataset is specifically designed for training text quality assessment classifiers, containing annotated…

Text Quality Classifier Training Dataset This dataset is specifically designed for training text quality assessment classifiers, containing annotated data from multiple high-quality corpora covering both English and Chinese texts across various professional domains. Dataset Summary Total Size: ~40B tokens (after sampling) Languages: English, Chinese Domains: General text, Mathematics, Programming, Reasoning & QA Annotation Dimensions: Mathematical intelligence… See the full description on the dataset page:

Source: Hugging Face Hub (OpenSQZ/Classifiers-Data). Metadata imported from the dataset’s Hub tags.

Advertisement