Benchmark

GroundUI-1K

multimodal vision multimodal

A subset of GroundUI-18K for UI grounding evaluation, where models must predict action coordinates on screenshots based on single-step instructions across web, desktop, and mobile platforms.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Nova Pro Amazon 81.4 来源 ↗
2 Nova Lite Amazon 80.2 来源 ↗