Benchmark
AGIEval
reasoning
general
math
text
A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission tests (LSAT), math competitions, lawyer qualification tests, and civil service exams. Contains 20 tasks (18 multiple-choice, 2 cloze) designed to assess understanding, knowledge, reasoning, and calculation abilities in real-world academic and professional contexts.
语言EN
满分1
参评模型5
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Mistral Small 3 24B Base | Mistral AI | 65.8 | 来源 ↗ |
| 2 | Gemma 2 27B | 55.1 | 来源 ↗ | |
| 3 | Gemma 2 9B | 52.8 | 来源 ↗ | |
| 4 | Granite 3.3 8B Base | IBM | 49.3 | 来源 ↗ |
| 5 | Ministral 8B Instruct | Mistral AI | 48.3 | 来源 ↗ |