Benchmark

BrowseComp

reasoning search text

BrowseComp is a benchmark comprising 1,266 questions that challenge AI agents to persistently navigate the internet in search of hard-to-find, entangled information. The benchmark measures agents' ability to exercise persistence in information gathering, demonstrate creativity in web navigation, and find concise, verifiable answers. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers.

语言EN
满分1
参评模型22

模型排名

名次 模型 机构 分数 来源
1 GPT-5.6 Sol OpenAI 90.4 来源 ↗
2 GPT-5.5 Pro OpenAI 90.1 来源 ↗
3 GPT-5.4 Pro OpenAI 89.3 来源 ↗
4 GPT-5.6 Terra OpenAI 87.5 来源 ↗
5 Claude Sonnet 5 Anthropic 84.7 来源 ↗
6 GPT-5.5 OpenAI 84.4 来源 ↗
7 MiniMax-M3 MiniMax 83.5 来源 ↗
8 GPT-5.6 Luna OpenAI 83.3 来源 ↗
9 GPT-5.4 OpenAI 82.7 来源 ↗
10 LongCat-2.0 Meituan 79.9 来源 ↗
11 Step 3.7 Flash StepFun 75.8 来源 ↗
12 GPT-5 OpenAI 54.9 来源 ↗
13 o4-mini OpenAI 51.5 来源 ↗
14 o3 OpenAI 49.7 来源 ↗
15 GLM-4.6 Zhipu AI 45.1 来源 ↗
16 Grok 4 Fast xAI 44.9 来源 ↗
17 Nemotron 3 Ultra 550B A55B NVIDIA 44.4 来源 ↗
18 DeepSeek-V3.2-Exp DeepSeek 40.1 来源 ↗
19 DeepSeek-V3.1 DeepSeek 30.0 来源 ↗
20 GLM-4.5 Zhipu AI 26.4 来源 ↗
21 GLM-4.5-Air Zhipu AI 21.3 来源 ↗
22 DeepSeek-R1-0528 DeepSeek 8.9 来源 ↗