Skip to content
Advertisement

Datasets

3,216 results

Text

FloresBitextMining

FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English…

1K–10K·CC-BY-SA·Parquet
MultimodalTabularText

panlex-meanings

Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…

100M–1B·CC0·CSV
Text

WMT18

WMT18

100M–1B·Custom / Research-only·Parquet
Text

Kreyòl-MT

Kreyòl-MT

1M–10M·Custom / Research-only·Parquet
Text

OPUS

Collection of OPUS Corpus from has been collected. The following corpora have been included: UNPC GlobalVoices TED2020 News-Commentary…

10M–100M·Parquet
Text

croissant_dataset

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…

>1B·Arrow
Text

IndicGenBenchFloresBitextMining

IndicGenBenchFloresBitextMining An MTEB dataset Massive Text Embedding Benchmark Flores-IN dataset is an extension of Flores dataset released as…

100K–1M·CC-BY-SA·Parquet
Text

WMT17

WMT17

10M–100M·Custom / Research-only·Parquet
Advertisement