Text Classification
109 results
sapnhap-bando-vn
sapnhap-bando-vn — Vietnam's 2025 administrative-merger atlas 🇻🇳 Tóm tắt. Một bản sao đầy đủ, có cấu trúc của…
thuvienphapluat.vn /tnpl/ Vietnamese Legal Terminology (bilingual VIEN)
thuvienphapluat.vn /tnpl/ Vietnamese Legal Terminology (bilingual VI EN)
SumTablets
SumTablets 🏺 A Transliteration Dataset of Sumerian Tablets Preprocessing scripts on GitHub. We welcome contributions! What is it?…
Turkish Academic Theses Abstracts (TR/EN)
Turkish Academic Theses Abstracts (TR/EN)
M4LE
Introduction M4LE is a Multi-ability, Multi-range, Multi-task, bilingual benchmark for long-context evaluation. We categorize long-context…
CodeMixBench
ℹ️Dataset Card for CodeMixBench EMNLP'25 CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic…
Maori_English_New_Zealand
Dataset for Translation from Maori to English The source of this dataset is scraped from the website TEARA.…
Mixed Arabic Datasets (MAD) Corpus
Mixed Arabic Datasets (MAD) Corpus
rtm-sgt-ocr-v1
Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…
JMedBench
Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any…
IPA Phonetic Lexicon (7M words)
IPA Phonetic Lexicon (7M words)
mmarco-contrastive
mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to…