Benchmark

Wild Bench

general reasoning text

WildBench is an automated evaluation framework that benchmarks large language models using 1,024 challenging, real-world tasks selected from over one million human-chatbot conversation logs. It introduces two evaluation metrics (WB-Reward and WB-Score) that achieve high correlation with human preferences and uses task-specific checklists for systematic evaluation.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Mistral Small 3.2 24B Instruct Mistral AI 65.3 来源 ↗
2 Mistral Small 3 24B Instruct Mistral AI 52.2 来源 ↗
3 Jamba 1.5 Large AI21 Labs 48.5 来源 ↗
4 Jamba 1.5 Mini AI21 Labs 42.4 来源 ↗