cc100-documents
cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…
575 results
cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…
Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for…
Open Food Facts Product Database
hhtools PARC MS — terrain-aware humanoid motion clips 中文说明 Motion clips in PARC MS layout for human-humanoid-tools (hhtools):…
MMLU-ProX Multilingual Model Predictions
Text Quality Classifier Training Dataset This dataset is specifically designed for training text quality assessment classifiers, containing annotated…
EU Law Dataset - Category 15.10
This is a re-upload of the aya collection, and only differs in the structure of upload. While the…
Dataset Summary GlotCC-V1.0 is a document-level, general domain dataset derived from CommonCrawl, covering more than 1000 languages.It is…
OpenAssistant Conversations Release 2
OpenAssistant Conversations
FollowCam Dataset Description FollowCam is a benchmark for evaluating video generation models on viewpoint transformation and subject following…