Benchmark
GroundUI-1K
multimodal
vision
multimodal
A subset of GroundUI-18K for UI grounding evaluation, where models must predict action coordinates on screenshots based on single-step instructions across web, desktop, and mobile platforms.
语言EN
满分1
参评模型2