Benchmark

ScreenSpot

vision multimodal spatial_reasoning multimodal

ScreenSpot is the first realistic GUI grounding benchmark that encompasses mobile, desktop, and web environments. The dataset comprises over 1,200 instructions from iOS, Android, macOS, Windows and Web environments, along with annotated element types (text and icon/widget), designed to evaluate visual GUI agents' ability to accurately locate screen elements based on natural language instructions.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 88.5 来源 ↗
2 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 87.1 来源 ↗
3 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 84.7 来源 ↗