Benchmark
HiddenMath
math
reasoning
text
Google DeepMind's internal mathematical reasoning benchmark that introduces novel problems not encountered during model training to evaluate true mathematical reasoning capabilities rather than memorization
语言EN
满分1
参评模型13
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemini 2.0 Flash | 63.0 | 来源 ↗ | |
| 2 | Gemma 3 27B | 60.3 | 来源 ↗ | |
| 3 | Gemini 2.0 Flash-Lite | 55.3 | 来源 ↗ | |
| 4 | Gemma 3 12B | 54.5 | 来源 ↗ | |
| 5 | Gemini 1.5 Pro | 52.0 | 来源 ↗ | |
| 6 | Gemini 1.5 Flash | 47.2 | 来源 ↗ | |
| 7 | Gemma 3 4B | 43.0 | 来源 ↗ | |
| 8 | Gemma 3n E4B Instructed | 37.7 | 来源 ↗ | |
| 9 | Gemma 3n E4B Instructed LiteRT Preview | 37.7 | 来源 ↗ | |
| 10 | Gemini 1.5 Flash 8B | 32.8 | 来源 ↗ | |
| 11 | Gemma 3n E2B Instructed | 27.7 | 来源 ↗ | |
| 12 | Gemma 3n E2B Instructed LiteRT (Preview) | 27.7 | 来源 ↗ | |
| 13 | Gemma 3 1B | 15.8 | 来源 ↗ |