Skip to content
Advertisement

10K–100K

748 results

Text

NusaTranslationBitextMining

NusaTranslationBitextMining An MTEB dataset Massive Text Embedding Benchmark NusaTranslation is a parallel dataset for machine translation on 11…

10K–100K·CC-BY-SA·Parquet
Text

SumTablets

SumTablets 🏺 A Transliteration Dataset of Sumerian Tablets Preprocessing scripts on GitHub. We welcome contributions! What is it?…

10K–100K·CC-BY·CSV
Text

translator-rules-dataset

Translator Rules Dataset A high-quality Arabic↔English machine-translation dataset built from the WAM (Web Archive Management) translation corpus —…

10K–100K·Parquet
MultimodalTabularText

M4LE

Introduction M4LE is a Multi-ability, Multi-range, Multi-task, bilingual benchmark for long-context evaluation. We categorize long-context…

10K–100K·MIT·JSON
Text

CodeMixBench

ℹ️Dataset Card for CodeMixBench EMNLP'25 CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic…

10K–100K·Apache-2.0·CSV
Text

ACES

ACES

10K–100K·CC-BY-NC-SA·JSON
Text

wmt24pp

WMT24++ This repository contains the human translation and post-edit data for the 55 en- xx language pairs released…

10K–100K·Apache-2.0·JSON
AudioMultimodalText

ScreenTalk_JA2ZH-XS

ScreenTalk JA2ZH-XS ScreenTalk JA2ZH-XS is a paired dataset of Japanese speech and Chinese translated text released by DataLabX.…

10K–100K·Parquet
MultimodalTabularText

TED_talks

!NOTE Dataset origin: Context TED is devoted to spreading powerful ideas in just about any topic. These datasets…

10K–100K·CSV
Text

ArzEn-MultiGenre

ArzEn-MultiGenre: A Comprehensive Parallel Dataset Overview ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection…

10K–100K·CC-BY·CSV
Text

ALMA-R-Preference

Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…

10K–100K·MIT·Parquet
Advertisement