Skip to content
Advertisement

Text

2,175 results

Text

croissant_dataset

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…

>1B·Arrow
Text

IndicGenBenchFloresBitextMining

IndicGenBenchFloresBitextMining An MTEB dataset Massive Text Embedding Benchmark Flores-IN dataset is an extension of Flores dataset released as…

100K–1M·CC-BY-SA·Parquet
Text

WMT17

WMT17

10M–100M·Custom / Research-only·Parquet
Text

WikidataLabels

Wikidata Labels Large parallel corpus for machine translation Entity label data extracted from Wikidata (2022-01-03), filtered for item…

100M–1B·CC0·Parquet
Text

WMT16

WMT16

1M–10M·Custom / Research-only·Parquet
Text

wmt24pp

WMT24++ This repository contains the human translation and post-edit data for the 55 en- xx language pairs released…

10K–100K·Apache-2.0·JSON
Text

BornholmBitextMining

BornholmBitextMining An MTEB dataset Massive Text Embedding Benchmark Danish Bornholmsk Parallel Corpus. Bornholmsk is a Danish dialect spoken…

1K–10K·CC-BY·Parquet
Text

WMT19

WMT19

100M–1B·Custom / Research-only·Parquet
Text

ulatroi

🧠 Project SLOB: Spontaneous Lifestyle & Observational Behaviors Dataset 📌 Abstract Welcome to the primary data ingestion node…

<1K·MIT·Text (raw)
Advertisement