Masked Language Modeling
11 results
Text
croissant_dataset_no_web_data
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Ressources are currently being uploaded ! Licenses Data redistributed here…
10M–100M·Arrow
Text
PubChem-124M-SMILES-SELFIES-InChI-IUPAC
PubChem-124M-Canonicalized-SELFIES-InChI-IUPAC Dataset Summary This dataset contains ~124 million chemical structures sourced from PubChem (as of Jan…
100M–1B·CC0
Text
croissant_dataset
CroissantLLM: A Truly Bilingual French-English Language Model Dataset Licenses Data redistributed here is subject to the original license…
>1B·Arrow
MultimodalTabularText
the-stack-mini-nonshuffled
This repo contains (up to) 30k samples of 21 languages (top 20 languages by StackOverflow survey, html/css was…
1M–10M·Custom / Research-only·Parquet
MultimodalTabularText
cc100-documents
cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…
100M–1B·Parquet