Benchmark

OCRBench

vision image-to-text multimodal

OCRBench: Comprehensive evaluation benchmark for assessing Optical Character Recognition (OCR) capabilities in Large Multimodal Models across text recognition, scene text VQA, and document understanding tasks

语言EN
满分1
参评模型7

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 VL 72B Instruct Alibaba Cloud / Qwen Team 88.5 来源 ↗
2 Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team 87.7 来源 ↗
3 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 86.4 来源 ↗
4 Phi-4-multimodal-instruct Microsoft 84.4 来源 ↗
5 DeepSeek VL2 Small DeepSeek 83.4 来源 ↗
6 DeepSeek VL2 DeepSeek 81.1 来源 ↗
7 DeepSeek VL2 Tiny DeepSeek 80.9 来源 ↗