Benchmark
C-Eval
general
reasoning
text
多语言
C-Eval is a comprehensive Chinese evaluation suite designed to assess advanced knowledge and reasoning abilities of foundation models in a Chinese context. It comprises 13,948 multiple-choice questions across 52 diverse disciplines spanning humanities, science, and engineering, with four difficulty levels: middle school, high school, college, and professional. The benchmark includes C-Eval Hard, a subset of very challenging subjects requiring advanced reasoning abilities.
语言EN
满分1
参评模型5
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Kimi K2 Base | Moonshot AI | 92.5 | 来源 ↗ |
| 2 | Kimi-k1.5 | Moonshot AI | 88.3 | 来源 ↗ |
| 3 | DeepSeek-V3 | DeepSeek | 86.5 | 来源 ↗ |
| 4 | Qwen2 72B Instruct | Alibaba Cloud / Qwen Team | 83.8 | 来源 ↗ |
| 5 | Qwen2 7B Instruct | Alibaba Cloud / Qwen Team | 77.2 | 来源 ↗ |