Benchmark

AlpacaEval 2.0

general creativity reasoning text

AlpacaEval 2.0 is a length-controlled automatic evaluator for instruction-following language models that uses GPT-4 Turbo to assess model responses against a baseline. It evaluates models on 805 diverse instruction-following tasks including creative writing, classification, programming, and general knowledge questions. The benchmark achieves 0.98 Spearman correlation with ChatBot Arena while being fast (< 3 minutes) and affordable (< $10 in OpenAI credits). It addresses length bias in automatic evaluation through length-controlled win-rates and uses weighted scoring based on response quality.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Granite 3.3 8B Base IBM 62.7 来源 ↗
2 Granite 3.3 8B Instruct IBM 62.7 来源 ↗
3 DeepSeek-V2.5 DeepSeek 50.5 来源 ↗
4 IBM Granite 4.0 Tiny Preview IBM 35.2 来源 ↗