Measuring Massive Multitask Language Understanding
Measuring Massive Multitask Language Understanding
1,233 results
Measuring Massive Multitask Language Understanding
GLUE (General Language Understanding Evaluation benchmark)
FineWeb Tokenized (AnisoleAI)
Colored MNIST Dataset A comprehensive dataset of MNIST digits with RGB colored backgrounds, designed for multi-objective classification tasks…
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing GOKU-2M is a large-scale, unified instruction-based…
Vero-600k Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with…
Synthetic Veterinary Ultrasound Dataset (AFAST)
KITTI Pseudo Depth (Eigen Split) with Depth Anything V2 Dataset Description This dataset provides high-quality pseudo depth maps…
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might…