Benchmark

AlignBench

general language math reasoning roleplay text 多语言

AlignBench is a comprehensive multi-dimensional benchmark for evaluating Chinese alignment of Large Language Models. It contains 8 main categories: Fundamental Language Ability, Advanced Chinese Understanding, Open-ended Questions, Writing Ability, Logical Reasoning, Mathematics, Task-oriented Role Play, and Professional Knowledge. The benchmark includes 683 real-scenario rooted queries with human-verified references and uses a rule-calibrated multi-dimensional LLM-as-Judge approach with Chain-of-Thought for evaluation.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen2.5 72B Instruct Alibaba Cloud / Qwen Team 81.6 来源 ↗
2 DeepSeek-V2.5 DeepSeek 80.4 来源 ↗
3 Qwen2.5 7B Instruct Alibaba Cloud / Qwen Team 73.3 来源 ↗
4 Qwen2 7B Instruct Alibaba Cloud / Qwen Team 72.1 来源 ↗