Benchmark

C-Eval

general reasoning text 多语言

C-Eval is a comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. It comprises 13,948 multiple-choice questions across 52 diverse disciplines spanning humanities, science, and engineering, with four difficulty levels: middle school, high school, college, and professional. The benchmark includes C-Eval Hard, a subset of very challenging subjects requiring advanced reasoning abilities.

语言EN
满分1
参评模型5

模型排名

名次 模型 机构 分数 来源
1 Kimi K2 Base Moonshot AI 92.5 来源 ↗
2 Kimi-k1.5 Moonshot AI 88.3 来源 ↗
3 DeepSeek-V3 DeepSeek 86.5 来源 ↗
4 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 83.8 来源 ↗
5 Qwen2 7B Instruct Alibaba Cloud / Qwen Team 77.2 来源 ↗