Benchmark

AGIEval

reasoning general math text

A human-centric benchmark for evaluating foundation models on standardized exams including college entrance exams (Gaokao, SAT), law school admission tests (LSAT), math competitions, lawyer qualification tests, and civil service exams. Contains 20 tasks (18 multiple-choice, 2 cloze) designed to assess understanding, knowledge, reasoning, and calculation abilities in real-world academic and professional contexts.

语言EN
满分1
参评模型5

模型排名

名次 模型 机构 分数 来源
1 Mistral Small 3 24B Base Mistral AI 65.8 来源 ↗
2 Gemma 2 27B Google 55.1 来源 ↗
3 Gemma 2 9B Google 52.8 来源 ↗
4 Granite 3.3 8B Base IBM 49.3 来源 ↗
5 Ministral 8B Instruct Mistral AI 48.3 来源 ↗