Benchmark

TextVQA

vision multimodal image-to-text multimodal

TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Introduced to benchmark VQA models' ability to read and reason about text within images, particularly for assistive technologies for visually impaired users. The dataset addresses the gap where existing VQA datasets had few text-based questions or were too small.

语言EN
满分1
参评模型15

模型排名

名次 模型 机构 分数 来源
1 Qwen2-VL-72B-Instruct Alibaba Cloud / Qwen Team 85.5 来源 ↗
2 Qwen2.5 VL 7B Instruct Alibaba Cloud / Qwen Team 84.9 来源 ↗
3 Qwen2.5-Omni-7B Alibaba Cloud / Qwen Team 84.4 来源 ↗
4 DeepSeek VL2 DeepSeek 84.2 来源 ↗
5 DeepSeek VL2 Small DeepSeek 83.4 来源 ↗
6 Nova Pro Amazon 81.5 来源 ↗
7 DeepSeek VL2 Tiny DeepSeek 80.7 来源 ↗
8 Nova Lite Amazon 80.2 来源 ↗
9 Grok-1.5V xAI 78.1 来源 ↗
10 Phi-4-multimodal-instruct Microsoft 75.6 来源 ↗
11 Llama 3.2 90B Instruct Meta 73.5 来源 ↗
12 Phi-3.5-vision-instruct Microsoft 72.0 来源 ↗
13 Gemma 3 12B Google 67.7 来源 ↗
14 Gemma 3 27B Google 65.1 来源 ↗
15 Gemma 3 4B Google 57.8 来源 ↗