quickmt-train.hu-en
quickmt hu-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…
253 results
quickmt hu-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…
20,000+ chinese sentences with translations and pinyin Source: Contributed by: Brian Vaughan Dataset Structure Each sample consists of:…
Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…
ScreenTalk JA2ZH-XS ScreenTalk JA2ZH-XS is a paired dataset of Japanese speech and Chinese translated text released by DataLabX.…
MyTown — US & Canadian local-government meetings
English-ASL Gloss Parallel Corpus 2012
English–Zomi Parallel Corpus (1.78M)
SEC 10-K Pre-processed
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…
!NOTE Dataset origin: Context TED is devoted to spreading powerful ideas in just about any topic. These datasets…
ChemrXiv Pdf Introducing ChemrXiv Pdf, a dataset that offers access to all PDFs published until September 15, 2024.…
Column Name Type Description 설명 Form str Registered word entry 단어 Part of Speech str or None Part…
ArzEn-MultiGenre: A Comprehensive Parallel Dataset Overview ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection…
Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…