Skip to content
Advertisement
ImageMultimodalText

VisualWebBench

VisualWebBench

VisualWebBench Dataset for the paper: VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? 🌐 Homepage 🐍 GitHub 📖 arXiv Introduction We introduce VisualWebBench, a multimodal benchmark designed to assess the understanding and grounding capabilities of MLLMs in web scenarios. VisualWebBench consists of seven tasks, and comprises 1.5K human-curated instances from 139 real websites, covering 87 sub-domains. We evaluate 14… See the full description on the dataset page:

Source: Hugging Face Hub (visualwebbench/VisualWebBench). Metadata imported from the dataset’s Hub tags.

Advertisement