Benchmark

Vibe-Eval

multimodal vision general multimodal

VIBE-Eval is a hard evaluation suite for measuring progress of multimodal language models, consisting of 269 visual understanding prompts with gold-standard responses authored by experts. The benchmark has dual objectives: vibe checking multimodal chat models for day-to-day tasks and rigorously testing frontier models, with the hard set containing >50% questions that all frontier models answer incorrectly.

语言EN
满分1
参评模型8

模型排名

名次 模型 机构 分数 来源
1 Gemini 2.5 Pro Preview 06-05 Google 67.2 来源 ↗
2 Gemini 2.5 Pro Google 65.6 来源 ↗
3 Gemini 2.5 Flash Google 65.4 来源 ↗
4 Gemini 2.0 Flash Google 56.3 来源 ↗
5 Gemini 1.5 Pro Google 53.9 来源 ↗
6 Gemini 2.5 Flash-Lite Google 51.3 来源 ↗
7 Gemini 1.5 Flash Google 48.9 来源 ↗
8 Gemini 1.5 Flash 8B Google 40.9 来源 ↗