finetranslations-sample-ar-en
Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…
1,501 results
Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…
Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…
Bangla–English Synthetic Parallel Corpus
IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…
Dataset Card for "undl ru2en aligned" More Information needed
RosettIA Chanka Quechua — Judicial Parallel Data (redistributable subset)
Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…
Dataset Card for WenYanWen English Parallel Dataset Summary The WenYanWen English Parallel dataset is a multilingual parallel corpus…
mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to…
MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation 📘 EMNLP 2025 Khai Le-Duc , Tuyen Tran , Bach Phan…
quickmt fr-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication, and basic filtering with quickmt:…
African Languages Lab Multi-Open