Benchmark
VQAv2 (val)
vision
multimodal
language
reasoning
multimodal
VQAv2 is a balanced Visual Question Answering dataset containing open-ended questions about images that require understanding of vision, language, and commonsense knowledge to answer. VQAv2 addresses bias issues from the original VQA dataset by collecting complementary images such that every question is associated with similar images that result in different answers, forcing models to actually understand visual content rather than relying on language priors.
语言EN
满分1
参评模型3
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemma 3 12B | 71.6 | 来源 ↗ | |
| 2 | Gemma 3 27B | 71.0 | 来源 ↗ | |
| 3 | Gemma 3 4B | 62.4 | 来源 ↗ |