Benchmark

CSimpleQA

general language text 多语言

Chinese SimpleQA is the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions. It contains 3,000 high-quality questions spanning 6 major topics with 99 diverse subtopics, designed to assess Chinese factual knowledge across humanities, science, engineering, culture, and society.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 84.3 来源 ↗
2 Kimi K2 Instruct Moonshot AI 78.4 来源 ↗
3 Kimi K2 Base Moonshot AI 77.6 来源 ↗
4 DeepSeek-V3 DeepSeek 64.8 来源 ↗