Skip to content
Advertisement

Parquet

1,501 results

Text

quickmt-train.hu-en

quickmt hu-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…

100M–1B·Parquet
Text

vietnamese_health_dataset

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

100K–1M·MIT·Parquet
Text

chr_en

Dataset Card for ChrEn Dataset Summary ChrEn is a Cherokee-English parallel dataset to facilitate machine translation research between…

100K–1M·Custom / Research-only·Parquet
AudioMultimodalText

ScreenTalk_JA2ZH-XS

ScreenTalk JA2ZH-XS ScreenTalk JA2ZH-XS is a paired dataset of Japanese speech and Chinese translated text released by DataLabX.…

10K–100K·Parquet
Text

OPUS DGT

OPUS DGT

1M–10M·Custom / Research-only·Parquet
Text

ALMA-R-Preference

Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…

10K–100K·MIT·Parquet
MultimodalTabularText

planetarium

Dataset Card for Planetarium🪐 Planetarium🪐 is a dataset and benchmark for assessing LLMs in translating natural language descriptions…

100K–1M·CC-BY·Parquet
MultimodalTabularText

undl_zh2en_aligned

联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score…

10M–100M·MIT·Parquet
Advertisement