Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using…
Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is built using the GlotLID language identification and Ungoliant pipeline from CommonCrawl.We release our pipeline as open-source at List of Languages: See to get the list of splits available. Usage (Huggingface Hub — Recommended)… See the full description on the dataset page:
Source: Hugging Face Hub (cis-lmu/GlotCC-V1). Metadata imported from the dataset’s Hub tags.