Benchmark

BrowseComp Long Context 128k

reasoning search text

A challenging benchmark for evaluating web browsing agents' ability to persistently navigate the internet and find hard-to-locate, entangled information. Comprises 1,266 questions requiring strategic reasoning, creative search, and interpretation of retrieved content, with short and easily verifiable answers.

语言EN
满分1
参评模型1

模型排名

名次 模型 机构 分数 来源
1 GPT-5 OpenAI 90.0 来源 ↗