Benchmark
TextVQA
vision
multimodal
image-to-text
multimodal
TextVQA contains 45,336 questions on 28,408 images that require reasoning about text to answer. Introduced to benchmark VQA models' ability to read and reason about text within images, particularly for assistive technologies for visually impaired users. The dataset addresses the gap where existing VQA datasets had few text-based questions or were too small.
语言EN
满分1
参评模型15
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2-VL-72B-Instruct | Alibaba Cloud / Qwen Team | 85.5 | 来源 ↗ |
| 2 | Qwen2.5 VL 7B Instruct | Alibaba Cloud / Qwen Team | 84.9 | 来源 ↗ |
| 3 | Qwen2.5-Omni-7B | Alibaba Cloud / Qwen Team | 84.4 | 来源 ↗ |
| 4 | DeepSeek VL2 | DeepSeek | 84.2 | 来源 ↗ |
| 5 | DeepSeek VL2 Small | DeepSeek | 83.4 | 来源 ↗ |
| 6 | Nova Pro | Amazon | 81.5 | 来源 ↗ |
| 7 | DeepSeek VL2 Tiny | DeepSeek | 80.7 | 来源 ↗ |
| 8 | Nova Lite | Amazon | 80.2 | 来源 ↗ |
| 9 | Grok-1.5V | xAI | 78.1 | 来源 ↗ |
| 10 | Phi-4-multimodal-instruct | Microsoft | 75.6 | 来源 ↗ |
| 11 | Llama 3.2 90B Instruct | Meta | 73.5 | 来源 ↗ |
| 12 | Phi-3.5-vision-instruct | Microsoft | 72.0 | 来源 ↗ |
| 13 | Gemma 3 12B | 67.7 | 来源 ↗ | |
| 14 | Gemma 3 27B | 65.1 | 来源 ↗ | |
| 15 | Gemma 3 4B | 57.8 | 来源 ↗ |