Benchmark
Arena Hard
general
reasoning
creativity
text
Arena-Hard-Auto is an automatic evaluation benchmark for instruction-tuned LLMs consisting of 500 challenging real-world prompts curated by BenchBuilder. It includes open-ended software engineering problems, mathematical questions, and creative writing tasks. The benchmark uses LLM-as-a-Judge methodology with GPT-4.1 and Gemini-2.5 as automatic judges to approximate human preference. Arena-Hard achieves 98.6% correlation with human preference rankings and provides 3x higher separation of model performances compared to MT-Bench, making it highly effective for distinguishing between models of similar quality.
语言EN
满分1
参评模型21
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen3 235B A22B | Alibaba Cloud / Qwen Team | 95.6 | 来源 ↗ |
| 2 | Qwen3 32B | Alibaba Cloud / Qwen Team | 93.8 | 来源 ↗ |
| 3 | Qwen3 30B A3B | Alibaba Cloud / Qwen Team | 91.0 | 来源 ↗ |
| 4 | Llama-3.3 Nemotron Super 49B v1 | NVIDIA | 88.3 | 来源 ↗ |
| 5 | Mistral Small 3 24B Instruct | Mistral AI | 87.6 | 来源 ↗ |
| 6 | Qwen2.5 72B Instruct | Alibaba Cloud / Qwen Team | 81.2 | 来源 ↗ |
| 7 | Phi 4 Reasoning Plus | Microsoft | 79.0 | 来源 ↗ |
| 8 | DeepSeek-V2.5 | DeepSeek | 76.2 | 来源 ↗ |
| 9 | Phi 4 | Microsoft | 75.4 | 来源 ↗ |
| 10 | Phi 4 Reasoning | Microsoft | 73.3 | 来源 ↗ |
| 11 | Ministral 8B Instruct | Mistral AI | 70.9 | 来源 ↗ |
| 12 | Jamba 1.5 Large | AI21 Labs | 65.4 | 来源 ↗ |
| 13 | Granite 3.3 8B Instruct | IBM | 57.6 | 来源 ↗ |
| 14 | Granite 3.3 8B Base | IBM | 57.6 | 来源 ↗ |
| 15 | Qwen2.5 7B Instruct | Alibaba Cloud / Qwen Team | 52.0 | 来源 ↗ |
| 16 | Jamba 1.5 Mini | AI21 Labs | 46.1 | 来源 ↗ |
| 17 | Mistral Small 3.2 24B Instruct | Mistral AI | 43.1 | 来源 ↗ |
| 18 | Phi-3.5-MoE-instruct | Microsoft | 37.9 | 来源 ↗ |
| 19 | Phi-3.5-mini-instruct | Microsoft | 37.0 | 来源 ↗ |
| 20 | Phi 4 Mini | Microsoft | 32.8 | 来源 ↗ |
| 21 | IBM Granite 4.0 Tiny Preview | IBM | 26.7 | 来源 ↗ |