Benchmark
AlpacaEval 2.0
general
creativity
reasoning
text
AlpacaEval 2.0 is a length-controlled automatic evaluator for instruction-following language models that uses GPT-4 Turbo to assess model responses against a baseline. It evaluates models on 805 diverse instruction-following tasks including creative writing, classification, programming, and general knowledge questions. The benchmark achieves 0.98 Spearman correlation with ChatBot Arena while being fast (< 3 minutes) and affordable (< $10 in OpenAI credits). It addresses length bias in automatic evaluation through length-controlled win-rates and uses weighted scoring based on response quality.
语言EN
满分1
参评模型4
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Granite 3.3 8B Base | IBM | 62.7 | 来源 ↗ |
| 2 | Granite 3.3 8B Instruct | IBM | 62.7 | 来源 ↗ |
| 3 | DeepSeek-V2.5 | DeepSeek | 50.5 | 来源 ↗ |
| 4 | IBM Granite 4.0 Tiny Preview | IBM | 35.2 | 来源 ↗ |