Skip to content
Advertisement

Masked Language Modeling

11 results

Text

chr_en

Dataset Card for ChrEn Dataset Summary ChrEn is a Cherokee-English parallel dataset to facilitate machine translation research between…

100K–1M·Custom / Research-only·Parquet
Text

croissant_dataset_no_web_data

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…

10M–100M·Arrow
Text

croissant_dataset

CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…

>1B·Arrow
MultimodalTabularText

cc100-documents

cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…

100M–1B·Parquet
Text

wikipedia

Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is…

10M–100M·Mixed·Parquet
Advertisement