Benchmark
AndroidWorld_SR
general
multimodal
reasoning
multimodal
AndroidWorld Success Rate (SR) benchmark - A dynamic benchmarking environment for autonomous agents operating on Android devices. Evaluates agents on 116 programmatic tasks across 20 real-world Android apps using multimodal inputs (screen screenshots, accessibility trees, and natural language instructions). Measures success rate of agents completing tasks like sending messages, creating calendar events, and navigating mobile interfaces. Published at ICLR 2025. Best current performance: 30.6% success rate (M3A agent) vs 80.0% human performance.
语言EN
满分1
参评模型3
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5 VL 72B Instruct | Alibaba Cloud / Qwen Team | 35.0 | 来源 ↗ |
| 2 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 25.5 | 来源 ↗ |
| 3 | Qwen2.5 VL 32B Instruct | Alibaba Cloud / Qwen Team | 22.0 | 来源 ↗ |