Skip to content
Advertisement

Multilingual

575 results

Text

Maori_English_New_Zealand

Dataset for Translation from Maori to English The source of this dataset is scraped from the website TEARA.…

1K–10K·Other·CSV
Audio

StreamUni

The training dataset for the paper 'StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model' Model…

1K–10K·Apache-2.0·Audio (folder)
Text

quickmt-train.hu-en

quickmt hu-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication and basic filtering with quickmt:…

100M–1B·Parquet
Text

20k-en-zh-translation-pinyin-hsk

20,000+ chinese sentences with translations and pinyin Source: Contributed by: Brian Vaughan Dataset Structure Each sample consists of:…

100K–1M·Text (raw)
Text

wmt24pp

WMT24++ This repository contains the human translation and post-edit data for the 55 en- xx language pairs released…

10K–100K·Apache-2.0·JSON
Text

vietnamese_health_dataset

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

100K–1M·MIT·Parquet
Text

chr_en

Dataset Card for ChrEn Dataset Summary ChrEn is a Cherokee-English parallel dataset to facilitate machine translation research between…

100K–1M·Custom / Research-only·Parquet
AudioMultimodalText

ScreenTalk_JA2ZH-XS

ScreenTalk JA2ZH-XS ScreenTalk JA2ZH-XS is a paired dataset of Japanese speech and Chinese translated text released by DataLabX.…

10K–100K·Parquet
Text

croissant_dataset_no_web_data

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…

10M–100M·Arrow
MultimodalTabularText

TED_talks

!NOTE Dataset origin: Context TED is devoted to spreading powerful ideas in just about any topic. These datasets…

10K–100K·CSV
Text

UltraLink

multi-lingual, knowledge-grounded, multi-round dialogue dataset and model Summary • Construction Process • Paper • UltraLink-LM • Github Dataset…

1M–10M·MIT·JSON
Text

OPUS DGT

OPUS DGT

1M–10M·Custom / Research-only·Parquet
Text

ArzEn-MultiGenre

ArzEn-MultiGenre: A Comprehensive Parallel Dataset Overview ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection…

10K–100K·CC-BY·CSV
Text

ALMA-R-Preference

Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…

10K–100K·MIT·Parquet
Advertisement