Benchmark

SimpleQA

general reasoning text

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains 4,326 short, fact-seeking questions that are adversarially collected and designed to have single, indisputable answers. Questions cover diverse topics from science and technology to entertainment, and the benchmark also measures model calibration by evaluating whether models know what they know.

语言EN
满分1
参评模型26

模型排名

名次 模型 机构 分数 来源
1 DeepSeek-V3.2-Exp DeepSeek 97.1 来源 ↗
2 Grok 4 Fast xAI 95.0 来源 ↗
3 DeepSeek-V3.1 DeepSeek 93.4 来源 ↗
4 DeepSeek-R1-0528 DeepSeek 92.3 来源 ↗
5 GPT-4.5 OpenAI 62.5 来源 ↗
6 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 54.3 来源 ↗
7 Gemini 2.5 Pro Preview 06-05 Google 54.0 来源 ↗
8 Gemini 2.5 Pro Google 50.8 来源 ↗
9 o1 OpenAI 47.0 来源 ↗
10 o1-preview OpenAI 42.4 来源 ↗
11 GPT-4o OpenAI 38.2 来源 ↗
12 Kimi K2 Base Moonshot AI 35.3 来源 ↗
13 Kimi K2-Instruct-0905 Moonshot AI 31.0 来源 ↗
14 Kimi K2 Instruct Moonshot AI 31.0 来源 ↗
15 Gemini 2.5 Flash Google 26.9 来源 ↗
16 DeepSeek-V3 DeepSeek 24.9 来源 ↗
17 Gemini 2.0 Flash-Lite Google 21.7 来源 ↗
18 o3-mini OpenAI 15.0 来源 ↗
19 Mistral Small 3.2 24B Instruct Mistral AI 12.1 来源 ↗
20 Gemini 2.5 Flash-Lite Google 10.7 来源 ↗
21 Mistral Small 3.1 24B Instruct Mistral AI 10.4 来源 ↗
22 Gemma 3 27B Google 10.0 来源 ↗
23 Gemma 3 12B Google 6.3 来源 ↗
24 Gemma 3 4B Google 4.0 来源 ↗
25 Phi 4 Microsoft 3.0 来源 ↗
26 Gemma 3 1B Google 2.2 来源 ↗