Benchmark
BrowseComp-zh
reasoning
search
text
多语言
A high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web, consisting of 289 multi-hop questions spanning 11 diverse domains including Film & TV, Technology, Medicine, and History. Questions are reverse-engineered from short, objective, and easily verifiable answers, requiring sophisticated reasoning and information reconciliation beyond basic retrieval. The benchmark addresses linguistic, infrastructural, and censorship-related complexities in Chinese web environments.
语言ZH
满分1
参评模型3
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | DeepSeek-V3.1 | DeepSeek | 49.2 | 来源 ↗ |
| 2 | DeepSeek-V3.2-Exp | DeepSeek | 47.9 | 来源 ↗ |
| 3 | DeepSeek-R1-0528 | DeepSeek | 35.7 | 来源 ↗ |