Skip to content
Advertisement

Parquet

1,501 results

Text

quickmt-train.zh-en

quickmt zh-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…

10M–100M·Parquet
MultimodalTabularText

bhasha-sft

Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual…

10M–100M·Mixed·Parquet
Text

skr_trans_distill

Dataset Card for skr trans distill 本数据集由 SakuraLLM 大模型生成机器翻译结果,主要用于模型蒸馏训练。数据集包含日文原文及其对应的中文翻译,适用于日译中任务的模型训练与蒸馏。 Dataset Details Dataset Description…

10M–100M·MIT·Parquet
Text

mala-bilingual-translation-corpus

MaLA Corpus: Massive Language Adaptation Corpus This MaLA-LM/mala-bilingual-translation-corpus is the MaLA bilingual translation corpus, collected…

>1B·ODC-BY·Parquet
Text

WMT15

WMT15

10M–100M·Custom / Research-only·Parquet
Text

flores

FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English…

1K–10K·CC-BY-SA·Parquet
Advertisement