Skip to content
Advertisement

Multilingual

575 results

MultimodalTabularText

bhasha-sft

Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual…

10M–100M·Mixed·Parquet
Text

LightNovel5000

Light novels translated in Chinese - crawled from public websites that do not prohibit crawlers 脚盆轻小说汉化 - 从未禁止爬虫的公共网站爬取…

<1K·Other
Text

IndonesianNMT

This dataset is used on the paper "Replicable Benchmarking of Neural Machine Translation (NMT) on Low-Resource Local Languages…

10K–100K
Text

tanglish-tamil

Translation of Tanglish to tamil Source: karky.in To use python import datasets s = datasets.load dataset('Deepakvictor/tanglish-tamil') print(s)…

<1K·OpenRAIL·CSV
Text

skr_trans_distill

Dataset Card for skr trans distill 本数据集由 SakuraLLM 大模型生成机器翻译结果,主要用于模型蒸馏训练。数据集包含日文原文及其对应的中文翻译,适用于日译中任务的模型训练与蒸馏。 Dataset Details Dataset Description…

10M–100M·MIT·Parquet
Text

WMT15

WMT15

10M–100M·Custom / Research-only·Parquet
Text

flores

FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English…

1K–10K·CC-BY-SA·Parquet
ImageMultimodalText

COCO-35L

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

100K–1M·Parquet
Advertisement