Benchmark
Vibe-Eval
multimodal
vision
general
multimodal
VIBE-Eval is a hard evaluation suite for measuring progress of multimodal language models, consisting of 269 visual understanding prompts with gold-standard responses authored by experts. The benchmark has dual objectives: vibe checking multimodal chat models for day-to-day tasks and rigorously testing frontier models, with the hard set containing >50% questions that all frontier models answer incorrectly.
语言EN
满分1
参评模型8
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini 2.5 Pro Preview 06-05 | 67.2 | 来源 ↗ | |
| 2 | Gemini 2.5 Pro | 65.6 | 来源 ↗ | |
| 3 | Gemini 2.5 Flash | 65.4 | 来源 ↗ | |
| 4 | Gemini 2.0 Flash | 56.3 | 来源 ↗ | |
| 5 | Gemini 1.5 Pro | 53.9 | 来源 ↗ | |
| 6 | Gemini 2.5 Flash-Lite | 51.3 | 来源 ↗ | |
| 7 | Gemini 1.5 Flash | 48.9 | 来源 ↗ | |
| 8 | Gemini 1.5 Flash 8B | 40.9 | 来源 ↗ |