Datasets
3,216 results
FineFineWeb
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming…
Multilingual Speech Commands Dataset (15 Languages, Augmented)
Multilingual Speech Commands Dataset (15 Languages, Augmented)
SWE-bench_Verified
Dataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been…
Measuring Massive Multitask Language Understanding
Measuring Massive Multitask Language Understanding
GLUE (General Language Understanding Evaluation benchmark)
GLUE (General Language Understanding Evaluation benchmark)
github-code-2025-language-split
📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was…
FineWeb Tokenized (AnisoleAI)
FineWeb Tokenized (AnisoleAI)
OpenThoughts-1k-sample
!NOTE We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample…
COCO-Caption
Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…
Imagenet21K
NOTE: I have recaptioned all images here This dataset is the entire 21K ImageNet dataset with about 13…
arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu
Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu.…
EasyNegative
Negative Embedding This is a Negative Embedding trained with Counterfeit. Please use it in the "stable-diffusion-webuiembeddings" folder.It can…
colored_mnist_28
Colored MNIST Dataset A comprehensive dataset of MNIST digits with RGB colored backgrounds, designed for multi-objective classification tasks…