GLUE (General Language Understanding Evaluation benchmark)
GLUE (General Language Understanding Evaluation benchmark)
2,175 results
GLUE (General Language Understanding Evaluation benchmark)
📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was…
FineWeb Tokenized (AnisoleAI)
!NOTE We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample…
Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…
NOTE: I have recaptioned all images here This dataset is the entire 21K ImageNet dataset with about 13…
The FIP 1.0 Data Set: Highly Resolved Annotated Image Time Series of 4,000 Wheat Plots Grown in Six…
@misc{kembhavi2016diagram, title={A Diagram Is Worth A Dozen Images}, author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon…
ASEAN Geolocation
Vero-600k Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with…
KITTI Pseudo Depth (Eigen Split) with Depth Anything V2 Dataset Description This dataset provides high-quality pseudo depth maps…
BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage 💻 Code 📖 Paper 📖 arXiv…
Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced…
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…
GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…