Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper 📖 arXiv GitHub Dataset Details As shown in…
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper 📖 arXiv GitHub Dataset Details As shown in the figure below, existing benchmarks lack consideration of the vision dependency of evaluation samples and potential data leakage from LLMs’ and LVLMs’ training data. Therefore, we introduce MMStar: an elite vision-indispensible multi-modal benchmark, aiming to ensure each curated sample exhibits… See the full description on the dataset page:
Source: Hugging Face Hub (Lin-Chen/MMStar). Metadata imported from the dataset’s Hub tags.