English-Polish MetricX-filtered Parallel Sentences with Qwen3 Embeddings
English-Polish MetricX-filtered Parallel Sentences with Qwen3 Embeddings
2,175 results
English-Polish MetricX-filtered Parallel Sentences with Qwen3 Embeddings
联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score…
Mixed Arabic Datasets (MAD) Corpus
Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…
Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…
OpenSakura Lilith LN COT Dataset
Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…
Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…
Bangla–English Synthetic Parallel Corpus
Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…
IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…
Dataset Card for "undl ru2en aligned" More Information needed
Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository: