Benchmark

VQAv2 (val)

vision multimodal language reasoning multimodal

VQAv2 is a balanced Visual Question Answering dataset containing open-ended questions about images that require understanding of vision, language, and commonsense knowledge to answer. VQAv2 addresses bias issues from the original VQA dataset by collecting complementary images such that every question is associated with similar images that result in different answers, forcing models to actually understand visual content rather than relying on language priors.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Gemma 3 12B Google 71.6 来源 ↗
2 Gemma 3 27B Google 71.0 来源 ↗
3 Gemma 3 4B Google 62.4 来源 ↗