Legal
35 results
thuvienphapluat.vn /tnpl/ Vietnamese Legal Terminology (bilingual VIEN)
thuvienphapluat.vn /tnpl/ Vietnamese Legal Terminology (bilingual VI EN)
NorwegianCourtsBitextMining
NorwegianCourtsBitextMining An MTEB dataset Massive Text Embedding Benchmark Nynorsk and Bokmål parallel corpus from Norwegian courts. Norwegian…
RosettIA Chanka Quechua — Judicial Parallel Data (redistributable subset)
RosettIA Chanka Quechua — Judicial Parallel Data (redistributable subset)
IN22GenBitextMining
IN22GenBitextMining An MTEB dataset Massive Text Embedding Benchmark IN22-Gen is a n-way parallel general-purpose multi-domain benchmark dataset for…
Marathi OCR model training dataset
Marathi OCR model training dataset
Open PII Masking 500k Ai4Privacy Dataset
Open PII Masking 500k Ai4Privacy Dataset
SDG-30K
SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image…
VideoHallu
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos Zongxia Li , Xiyang Wu , Guangyao Shi, Yubin…
InfoRe Technology public dataset №1
InfoRe Technology public dataset №1
Korean Single Speaker Speech Dataset
Korean Single Speaker Speech Dataset
SaSLaW
This repository contains the data of SaSLaW corpus. You can download it via the following command: huggingface-cli download…