Benchmark
Global-MMLU-Lite
general
language
reasoning
text
多语言
A lightweight version of Global MMLU benchmark that evaluates language models across multiple languages while addressing cultural and linguistic biases in multilingual evaluation.
语言EN
满分1
参评模型14
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini 2.5 Pro Preview 06-05 | 89.2 | 来源 ↗ | |
| 2 | Gemini 2.5 Pro | 88.6 | 来源 ↗ | |
| 3 | Gemini 2.5 Flash | 88.4 | 来源 ↗ | |
| 4 | Gemini 2.5 Flash-Lite | 81.1 | 来源 ↗ | |
| 5 | Gemini 2.0 Flash-Lite | 78.2 | 来源 ↗ | |
| 6 | Gemma 3 27B | 75.1 | 来源 ↗ | |
| 7 | Gemma 3 12B | 69.5 | 来源 ↗ | |
| 8 | Gemini Diffusion | 69.1 | 来源 ↗ | |
| 9 | Gemma 3n E4B Instructed | 64.5 | 来源 ↗ | |
| 10 | Gemma 3n E4B Instructed LiteRT Preview | 64.5 | 来源 ↗ | |
| 11 | Gemma 3n E2B Instructed | 59.0 | 来源 ↗ | |
| 12 | Gemma 3n E2B Instructed LiteRT (Preview) | 59.0 | 来源 ↗ | |
| 13 | Gemma 3 4B | 54.5 | 来源 ↗ | |
| 14 | Gemma 3 1B | 34.2 | 来源 ↗ |