Benchmark
AlignBench
general
language
math
reasoning
roleplay
text
多语言
AlignBench is a comprehensive multi-dimensional benchmark for evaluating Chinese alignment of Large Language Models. It contains 8 main categories: Fundamental Language Ability, Advanced Chinese Understanding, Open-ended Questions, Writing Ability, Logical Reasoning, Mathematics, Task-oriented Role Play, and Professional Knowledge. The benchmark includes 683 real-scenario rooted queries with human-verified references and uses a rule-calibrated multi-dimensional LLM-as-Judge approach with Chain-of-Thought for evaluation.
语言EN
满分1
参评模型4
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen2.5 72B Instruct | Alibaba Cloud / Qwen Team | 81.6 | 来源 ↗ |
| 2 | DeepSeek-V2.5 | DeepSeek | 80.4 | 来源 ↗ |
| 3 | Qwen2.5 7B Instruct | Alibaba Cloud / Qwen Team | 73.3 | 来源 ↗ |
| 4 | Qwen2 7B Instruct | Alibaba Cloud / Qwen Team | 72.1 | 来源 ↗ |