Skip to content
Advertisement

Parquet

1,501 results

Text

finetranslations-sample-ar-en

Sample of filtered and split into sentences by this script intended to be used for training sentence-level machine…

100M–1B·CC0·Parquet
Text

GH_text2code

Docstring to code data Dataset Summary This dataset contains pairs of English text and code from multiple programming…

10M–100M·Parquet
Text

IN22ConvBitextMining

IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…

1K–10K·CC-BY·Parquet
Text

OpusGnome

OpusGnome

10K–100K·Custom / Research-only·Parquet
ImageMultimodalText

Flickr30k

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

10K–100K·CC-BY·Parquet
Text

UnGa

UnGa

1M–10M·Custom / Research-only·Parquet
Text

OPUS-100

OPUS-100

10M–100M·Custom / Research-only·Parquet
Text

WenYanWen_English_Parallel

Dataset Card for WenYanWen English Parallel Dataset Summary The WenYanWen English Parallel dataset is a multilingual parallel corpus…

1M–10M·MIT·Parquet
Text

mmarco-contrastive

mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to…

100K–1M·Apache-2.0·Parquet
AudioMultimodalText

MultiMed-ST

MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech Translation 📘 EMNLP 2025 Khai Le-Duc , Tuyen Tran , Bach Phan…

10K–100K·MIT·Parquet
Text

quickmt-train.fr-en

quickmt fr-en Training Corpus Contains the following datasets downloaded with mtdata after deduplication, and basic filtering with quickmt:…

100M–1B·Parquet
Advertisement