Benchmark

VisualWebBench

vision multimodal frontend_development multimodal

A multimodal benchmark designed to assess the capabilities of multimodal large language models (MLLMs) across web page understanding and grounding tasks. Comprises 7 tasks (captioning, webpage QA, heading OCR, element OCR, element grounding, action prediction, and action grounding) with 1.5K human-curated instances from 139 real websites across 87 sub-domains.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Nova Pro Amazon 79.7 来源 ↗
2 Nova Lite Amazon 77.7 来源 ↗