Skip to content
Advertisement
Text

translator-rules-dataset

Translator Rules Dataset A high-quality Arabic↔English machine-translation dataset built from the WAM (Web Archive Management) translation corpus —…

Translator Rules Dataset A high-quality Arabic↔English machine-translation dataset built from the WAM (Web Archive Management) translation corpus — 10 years of professional Arabic↔English news/government translations — augmented with 200 LLM-extracted translation rules embedded in system prompts, and mixed with ~7% general-instruction data to prevent catastrophic forgetting during fine-tuning. Dataset Statistics Split Rows MT rows General rows General %… See the full description on the dataset page:

Source: Hugging Face Hub (sevka-k1/translator-rules-dataset). Metadata imported from the dataset’s Hub tags.

Advertisement