Benchmark

HLE

reasoning math multimodal

Humanity's Last Exam (HLE) is a multi-modal academic benchmark with 2,500 questions across mathematics, humanities, and natural sciences, designed to test LLM capabilities at the frontier of human knowledge with unambiguous, verifiable solutions

语言EN
满分1
参评模型5

模型排名

名次 模型 机构 分数 来源
1 Grok 4 Fast xAI 20.0 来源 ↗
2 GLM-4.6 Zhipu AI 17.2 来源 ↗
3 GLM-4.5 Zhipu AI 14.4 来源 ↗
4 GLM-4.5-Air Zhipu AI 10.6 来源 ↗
5 Kimi K2-Instruct-0905 Moonshot AI 4.7 来源 ↗