Skip to content
Advertisement

Text Classification

109 results

ImageMultimodalTabular

sapnhap-bando-vn

sapnhap-bando-vn — Vietnam's 2025 administrative-merger atlas 🇻🇳 Tóm tắt. Một bản sao đầy đủ, có cấu trúc của…

10K–100K·CC-BY-NC·Parquet
Text

aaaAAA

Writing-ai Crafting Historical Romance Novels with Deep Psychological Profiling in Lesbian Relationships

<1K·JSON
Text

SumTablets

SumTablets 🏺 A Transliteration Dataset of Sumerian Tablets Preprocessing scripts on GitHub. We welcome contributions! What is it?…

10K–100K·CC-BY·CSV
MultimodalTabularText

M4LE

Introduction M4LE is a Multi-ability, Multi-range, Multi-task, bilingual benchmark for long-context evaluation. We categorize long-context…

10K–100K·MIT·JSON
Text

CodeMixBench

ℹ️Dataset Card for CodeMixBench EMNLP'25 CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic…

10K–100K·Apache-2.0·CSV
Text

Maori_English_New_Zealand

Dataset for Translation from Maori to English The source of this dataset is scraped from the website TEARA.…

1K–10K·Other·CSV
Text

rtm-sgt-ocr-v1

Data Introduction Over 1.5 Million synthetically generated ground-truth/OCR pairs for post correction tasks from our paper "Large Synthetic…

1M–10M·Apache-2.0·CSV
MultimodalTabularText

JMedBench

Maintainers Junfeng Jiang@Aizawa Lab: jiangjf (at) is.s.u-tokyo.ac.jp Jiahao Huang@Aizawa Lab: jiahao-huang (at) g.ecc.u-tokyo.ac.jp If you find any…

100K–1M·JSON
Text

mmarco-contrastive

mMARCO-contrastive The dataset is a modification of mMARCO focusing on French and English parts. The aim is to…

100K–1M·Apache-2.0·Parquet
Advertisement