Skip to content
Advertisement

MIT

552 results

MultimodalTabularText

urbex

Urbex — Global Abandoned Locations Dataset (100k+) Safety and Legal Notice: Urban exploration may constitute trespass and can…

100K–1M·MIT·Parquet
MultimodalTabularText

M4LE

Introduction M4LE is a Multi-ability, Multi-range, Multi-task, bilingual benchmark for long-context evaluation. We categorize long-context…

10K–100K·MIT·JSON
Text

vietnamese_health_dataset

Team and Homepage Official Website: Hugging Face Organization: Contact If you encounter any issues with the dataset or…

100K–1M·MIT·Parquet
Text

UltraLink

multi-lingual, knowledge-grounded, multi-round dialogue dataset and model Summary • Construction Process • Paper • UltraLink-LM • Github Dataset…

1M–10M·MIT·JSON
Text

ALMA-R-Preference

Dataset Card for "ALMA-R-Preference" This is triplet preference data used by ALMA-R model. The triplet preference data, supporting…

10K–100K·MIT·Parquet
MultimodalTabularText

undl_zh2en_aligned

联合国数字图书馆的段落级中-英对齐平行语料 用我口胡的方法弄出来的平行语料,统计数据和拿argostranslate直接又跑了一份bleu score的结果已经丢论文里了,论文在写了在写了。应该拿这份去练机翻模型没问题,数据源是人写的。 bleu score…

10M–100M·MIT·Parquet
Text

Translate_datasets

Datasets From convert to RWKV datasets format It is recommended to shuffle before use Chinese - English &…

10M–100M·MIT·Arrow
Text

Parallel_Dataset

Dataset Sources Paper: LegoMT2: Selective Asynchronous Sharded Data Parallel Training for Massive Neural Machine Translation Link: Repository:

1M–10M·MIT·JSON
Text

NL2SH-ALFA

Dataset Card for NL2SH-ALFA This dataset is a collection of natural language (English) instructions and corresponding Bash commands…

10K–100K·MIT·CSV
ImageMultimodalTabular

ChEBI-20-MM

ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is…

10K–100K·MIT·CSV
Text

WenYanWen_English_Parallel

Dataset Card for WenYanWen English Parallel Dataset Summary The WenYanWen English Parallel dataset is a multilingual parallel corpus…

1M–10M·MIT·Parquet
Advertisement