Skip to content
Advertisement

Open

2,922 results

Text

WMT15

WMT15

10M–100M·Custom / Research-only·Parquet
Text

flores

FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English…

1K–10K·CC-BY-SA·Parquet
AudioMultimodalText

wmt-human-all-TTS

WMT Human + TTS Audio WMT human evaluation data (zouharvi/wmt-human-all) extended with TTS-synthesised source audio, covering 49 language…

100K–1M·CC-BY·Parquet
ImageMultimodalText

COCO-35L

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

100K–1M·Parquet
Text

IN22GenBitextMining

IN22GenBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Gen is a n-way parallel general-purpose multi-domain benchmark dataset for…

100K–1M·CC-BY·Parquet
ImageMultimodalText

CC3M-35L

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

1M–10M·Parquet
Text

TinyStories-Multilingual

Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short,…

10K–100K·Apache-2.0·JSON
Advertisement