Open
2,922 results
Mixed Arabic Datasets (MAD) Corpus
Mixed Arabic Datasets (MAD) Corpus
Translate_datasets
Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…
rtm-sgt-ocr-v1
Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…
OpenSakura Lilith LN COT Dataset
OpenSakura Lilith LN COT Dataset
finetranslations-sample-ar-en
Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…
GH_text2code
Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…
Bangla–English Synthetic Parallel Corpus
Bangla–English Synthetic Parallel Corpus
panlex-meanings
Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…
IN22ConvBitextMining
IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…
undl_ru2en_aligned
Dataset Card for "undl ru2en aligned" More Information needed
Parallel_Dataset
Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:
NL2SH-ALFA
Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…