Benchmark

BrowseComp-zh

reasoning search text 多语言

A high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web, consisting of 289 multi-hop questions spanning 11 diverse domains including Film & TV, Technology, Medicine, and History. Questions are reverse-engineered from short, objective, and easily verifiable answers, requiring sophisticated reasoning and information reconciliation beyond basic retrieval. The benchmark addresses linguistic, infrastructural, and censorship-related complexities in Chinese web environments.

语言ZH
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 DeepSeek-V3.1 DeepSeek 49.2 来源 ↗
2 DeepSeek-V3.2-Exp DeepSeek 47.9 来源 ↗
3 DeepSeek-R1-0528 DeepSeek 35.7 来源 ↗