Benchmark
BrowseComp
reasoning
search
text
BrowseComp is a benchmark comprising 1,266 questions that challenge AI agents to persistently navigate the internet in search of hard-to-find, entangled information. The benchmark measures agents' ability to exercise persistence in information gathering, demonstrate creativity in web navigation, and find concise, verifiable answers. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers.
语言EN
满分1
参评模型22
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | 90.4 | 来源 ↗ |
| 2 | GPT-5.5 Pro | OpenAI | 90.1 | 来源 ↗ |
| 3 | GPT-5.4 Pro | OpenAI | 89.3 | 来源 ↗ |
| 4 | GPT-5.6 Terra | OpenAI | 87.5 | 来源 ↗ |
| 5 | Claude Sonnet 5 | Anthropic | 84.7 | 来源 ↗ |
| 6 | GPT-5.5 | OpenAI | 84.4 | 来源 ↗ |
| 7 | MiniMax-M3 | MiniMax | 83.5 | 来源 ↗ |
| 8 | GPT-5.6 Luna | OpenAI | 83.3 | 来源 ↗ |
| 9 | GPT-5.4 | OpenAI | 82.7 | 来源 ↗ |
| 10 | LongCat-2.0 | Meituan | 79.9 | 来源 ↗ |
| 11 | Step 3.7 Flash | StepFun | 75.8 | 来源 ↗ |
| 12 | GPT-5 | OpenAI | 54.9 | 来源 ↗ |
| 13 | o4-mini | OpenAI | 51.5 | 来源 ↗ |
| 14 | o3 | OpenAI | 49.7 | 来源 ↗ |
| 15 | GLM-4.6 | Zhipu AI | 45.1 | 来源 ↗ |
| 16 | Grok 4 Fast | xAI | 44.9 | 来源 ↗ |
| 17 | Nemotron 3 Ultra 550B A55B | NVIDIA | 44.4 | 来源 ↗ |
| 18 | DeepSeek-V3.2-Exp | DeepSeek | 40.1 | 来源 ↗ |
| 19 | DeepSeek-V3.1 | DeepSeek | 30.0 | 来源 ↗ |
| 20 | GLM-4.5 | Zhipu AI | 26.4 | 来源 ↗ |
| 21 | GLM-4.5-Air | Zhipu AI | 21.3 | 来源 ↗ |
| 22 | DeepSeek-R1-0528 | DeepSeek | 8.9 | 来源 ↗ |