Skip to content
Advertisement
MultimodalTabularText

GlotCC-V1

Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using…

Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at List of Languages: See to get the list of splits available. Usage (Huggingface Hub — Recommended)… See the full description on the dataset page:

Source: Hugging Face Hub (cis-lmu/GlotCC-V1). Metadata imported from the dataset’s Hub tags.

Advertisement