skr_trans_distill
Dataset Card for skr trans distill 本数据集由 SakuraLLM 大模型生成机器翻译结果,主要用于模型蒸馏训练。数据集包含日文原文及其对应的中文翻译,适用于日译中任务的模型训练与蒸馏。 Dataset Details Dataset Description…
2,175 results
Dataset Card for skr trans distill 本数据集由 SakuraLLM 大模型生成机器翻译结果,主要用于模型蒸馏训练。数据集包含日文原文及其对应的中文翻译,适用于日译中任务的模型训练与蒸馏。 Dataset Details Dataset Description…
MaLA Corpus: Massive Language Adaptation Corpus This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected…
Synthetic Parallel EN-LG (external)
LAION-COCO translated to 200 languages
TED Translation Decision Dataset
WMT20 - MultiLingual Quality Estimation (MLQE) Task2
SETimes – A Parallel Corpus of English and South-East European Languages
WMT Human + TTS Audio WMT human evaluation data (zouharvi/wmt-human-all) extended with TTS-synthesised source audio, covering 49 language…
Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…
Shaistagi (شائستگی) Clean Urdu Mega-Dataset
IN22GenBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Gen is a n-way parallel general-purpose multi-domain benchmark dataset for…
CLERC Épée v0.2 — Sign Language Data Layer
Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…
Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short,…