Skip to content
Advertisement

CC-BY

672 results

Text

IN22ConvBitextMining

IN22ConvBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Conv is a n-way parallel conversation domain benchmark dataset for…

1K–10K·CC-BY·Parquet
ImageMultimodalText

Flickr30k

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

10K–100K·CC-BY·Parquet
AudioMultimodalText

wmt-human-all-TTS

WMT Human + TTS Audio WMT human evaluation data (zouharvi/wmt-human-all) extended with TTS-synthesised source audio, covering 49 language…

100K–1M·CC-BY·Parquet
Text

IN22GenBitextMining

IN22GenBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Gen is a n-way parallel general-purpose multi-domain benchmark dataset for…

100K–1M·CC-BY·Parquet
Text

tatoeba-bitext-mining

Tatoeba An MTEB dataset Massive Text Embedding Benchmark 1,000 English-aligned sentence pairs for each language based on the…

100K–1M·CC-BY·JSON
Text

BornholmBitextMining

BornholmBitextMining An MTEB dataset Massive Text Embedding Benchmark Danish Bornholmsk Parallel Corpus. Bornholmsk is a Danish dialect spoken…

1K–10K·CC-BY·Parquet
Text

PHINC

Abstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities,…

10K–100K·CC-BY·CSV
Advertisement