Datasets
3,216 results
Text
FloresBitextMining
FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English…
1K–10K·CC-BY-SA·Parquet
MultimodalTabularText
panlex-meanings
Dataset Card for panlex-meanings This is a dataset of words in several thousand languages, extracted from Dataset Details…
100M–1B·CC0·CSV
MultimodalTabularText
Mixed Arabic Datasets (MAD) Corpus
Mixed Arabic Datasets (MAD) Corpus
100M–1B·Parquet
Text
Alexandria Multudialectal Arabic Conversational Dataset for Machine Translation
Alexandria Multudialectal Arabic Conversational Dataset for Machine Translation
10K–100K·CC-BY-NC·Parquet
Text
croissant_dataset
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…
>1B·Arrow
Text
IndicGenBenchFloresBitextMining
IndicGenBenchFloresBitextMining An MTEB dataset Massive Text Embedding Benchmark Flores-IN dataset is an extension of Flores dataset released as…
100K–1M·CC-BY-SA·Parquet
MultimodalTabularText
OpenSakura Eve LN Aligned Dataset
OpenSakura Eve LN Aligned Dataset
100K–1M·Custom / Research-only·Parquet