Mixed Arabic Datasets (MAD) Corpus
Mixed Arabic Datasets (MAD) Corpus
2,594 results
Mixed Arabic Datasets (MAD) Corpus
Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…
Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…
OpenSakura Lilith LN COT Dataset
Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…
Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…
Bangla–English Synthetic Parallel Corpus
Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…
IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…
Dataset Card for "undl ru2en aligned" More Information needed
Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:
Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…
ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is…
IPA Phonetic Lexicon (7M words)