Benchmark
HLE
reasoning
math
multimodal
Humanity's Last Exam (HLE) is a multi-modal academic benchmark with 2,500 questions across mathematics, humanities, and natural sciences, designed to test LLM capabilities at the frontier of human knowledge with unambiguous, verifiable solutions
语言EN
满分1
参评模型5
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Grok 4 Fast | xAI | 20.0 | 来源 ↗ |
| 2 | GLM-4.6 | Zhipu AI | 17.2 | 来源 ↗ |
| 3 | GLM-4.5 | Zhipu AI | 14.4 | 来源 ↗ |
| 4 | GLM-4.5-Air | Zhipu AI | 10.6 | 来源 ↗ |
| 5 | Kimi K2-Instruct-0905 | Moonshot AI | 4.7 | 来源 ↗ |