Benchmark

SuperGPQA

reasoning general math legal healthcare finance chemistry economics physics text

SuperGPQA is a comprehensive benchmark that evaluates large language models across 285 graduate-level academic disciplines. The benchmark contains 25,957 questions covering 13 broad disciplinary areas including Engineering, Medicine, Science, and Law, with specialized fields in light industry, agriculture, and service-oriented domains. It employs a Human-LLM collaborative filtering mechanism with over 80 expert annotators to create challenging questions that assess graduate-level knowledge and reasoning capabilities.

语言EN
满分1
参评模型8

模型排名

名次 模型 机构 分数 来源
1 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 64.9 来源 ↗
2 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 62.6 来源 ↗
3 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 60.8 来源 ↗
4 Qwen3-Next-80B-A3B-Instruct Alibaba Cloud / Qwen Team 58.8 来源 ↗
5 Kimi K2 Instruct Moonshot AI 57.2 来源 ↗
6 Kimi K2-Instruct-0905 Moonshot AI 57.2 来源 ↗
7 Kimi K2 Base Moonshot AI 44.7 来源 ↗
8 Qwen3 235B A22B Alibaba Cloud / Qwen Team 44.1 来源 ↗