Benchmark

OSWorld

multimodal general vision multimodal

OSWorld: The first-of-its-kind scalable, real computer environment for multimodal agents, supporting task setup, execution-based evaluation, and interactive learning across Ubuntu, Windows, and macOS with 369 computer tasks involving real web and desktop applications, OS file I/O, and multi-application workflows

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 GPT-5.6 Sol OpenAI 62.6 来源 ↗
2 Claude Sonnet 4.5 Anthropic 61.4 来源 ↗
3 Claude Haiku 4.5 Anthropic 50.7 来源 ↗
4 GPT-5.6 Terra OpenAI 50.2 来源 ↗
5 GPT-5.6 Luna OpenAI 45.6 来源 ↗
6 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 8.8 来源 ↗
7 Qwen2.5 VL 32B Instruct Alibaba Cloud / Qwen Team 5.9 来源 ↗