Datasets
3,216 results
NIKL-korean-english-dictionary
Column Name Type Description 설명 Form str Registered word entry 단어 Part of Speech str or None Part…
ArzEn-MultiGenre
ArzEn-MultiGenre: A Comprehensive Parallel Dataset Overview ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection…
ALMA-R-Preference
Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…
planetarium
Dataset Card for Planetarium🪐 Planetarium🪐 is a dataset and benchmark for assessing LLMs in translating natural language descriptions…
English-Polish MetricX-filtered Parallel Sentences with Qwen3 Embeddings
English-Polish MetricX-filtered Parallel Sentences with Qwen3 Embeddings
undl_zh2en_aligned
联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score…
Mixed Arabic Datasets (MAD) Corpus
Mixed Arabic Datasets (MAD) Corpus
Translate_datasets
Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…
rtm-sgt-ocr-v1
Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…
OpenSakura Lilith LN COT Dataset
OpenSakura Lilith LN COT Dataset
finetranslations-sample-ar-en
Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…
GH_text2code
Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…