Skip to content
Advertisement

Translation

253 results

Text

wmt24pp

WMT24++ This repository contains the human translation and post-edit data for the 55 en- xx language pairs released…

10K–100K·Apache-2.0·JSON
Text

BornholmBitextMining

BornholmBitextMining An MTEB dataset Massive Text Embedding Benchmark Danish Bornholmsk Parallel Corpus. Bornholmsk is a Danish dialect spoken…

1K–10K·CC-BY·Parquet
Text

WMT19

WMT19

100M–1B·Custom / Research-only·Parquet
Text

ulatroi

🧠 Project SLOB: Spontaneous Lifestyle & Observational Behaviors Dataset 📌 Abstract Welcome to the primary data ingestion node…

<1K·MIT·Text (raw)
Text

OPUS-100

OPUS-100

10M–100M·Custom / Research-only·Parquet
Text

WMT14

WMT14

10M–100M·Custom / Research-only·Parquet
Text

OpusBooks

OpusBooks

1M–10M·Custom / Research-only·Parquet
Text

PHINC

Abstract Code-mixing is the phenomenon of using more than one language in a sentence. In the multilingual communities,…

10K–100K·CC-BY·CSV
Text

multi30k

Multi30k This dataset contains the "multi30k" dataset, which is the "task 1" dataset from here. Each example consists…

10K–100K·JSON
AudioMultimodalText

dowis

Do What I Say (DOWIS): A Spoken Prompt Dataset for Instruction-Following NEW DOWIS now also contains spoken and…

1K–10K·CC-BY·Parquet
MultimodalTabularText

UniST

UniST This dataset contains UniST codec-token training data exported from local metadata and codec results. We train UniSS…

10M–100M·CC-BY-NC
Advertisement