Skip to content
Advertisement

Datasets

3,216 results

Text

UltraLink

multi-lingual, knowledge-grounded, multi-round dialogue dataset and model Summary • Construction Process • Paper • UltraLink-LM • Github Dataset…

1M–10M·MIT·JSON
Text

OPUS DGT

OPUS DGT

1M–10M·Custom / Research-only·Parquet
Text

ArzEn-MultiGenre

ArzEn-MultiGenre: A Comprehensive Parallel Dataset Overview ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection…

10K–100K·CC-BY·CSV
Text

ALMA-R-Preference

Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…

10K–100K·MIT·Parquet
MultimodalTabularText

planetarium

Dataset Card for Planetarium🪐 Planetarium🪐 is a dataset and benchmark for assessing LLMs in translating natural language descriptions…

100K–1M·CC-BY·Parquet
MultimodalTabularText

undl_zh2en_aligned

联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score…

10M–100M·MIT·Parquet
Text

Translate_datasets

Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…

10M–100M·MIT·Arrow
Text

rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…

1M–10M·Apache-2.0·CSV
Text

finetranslations-sample-ar-en

Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…

100M–1B·CC0·Parquet
Text

GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…

10M–100M·Parquet
Advertisement