Benchmark

RealWorldQA

vision spatial_reasoning multimodal

RealWorldQA is a benchmark designed to evaluate basic real-world spatial understanding capabilities of multimodal models. The initial release consists of over 700 anonymized images taken from vehicles and other real-world scenarios, each accompanied by a question and easily verifiable answer. Released by xAI as part of their Grok-1.5 Vision preview to test models' ability to understand natural scenes and spatial relationships in everyday visual contexts.

语言EN
满分1
参评模型6

模型排名

名次 模型 机构 分数 来源
1 Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team 77.8 来源 ↗
2 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 70.3 来源 ↗
3 Grok-1.5V xAI 68.7 来源 ↗
4 DeepSeek VL2 DeepSeek 68.4 来源 ↗
5 DeepSeek VL2 Small DeepSeek 65.4 来源 ↗
6 DeepSeek VL2 Tiny DeepSeek 64.2 来源 ↗