Benchmark

Arena Hard

general reasoning creativity text

Arena-Hard-Auto is an automatic evaluation benchmark for instruction-tuned LLMs consisting of 500 challenging real-world prompts curated by BenchBuilder. It includes open-ended software engineering problems, mathematical questions, and creative writing tasks. The benchmark uses LLM-as-a-Judge methodology with GPT-4.1 and Gemini-2.5 as automatic judges to approximate human preference. Arena-Hard achieves 98.6% correlation with human preference rankings and provides 3x higher separation of model performances compared to MT-Bench, making it highly effective for distinguishing between models of similar quality.

语言EN
满分1
参评模型21

模型排名

名次 模型 机构 分数 来源
1 Qwen3 235B A22B Alibaba Cloud / Qwen Team 95.6 来源 ↗
2 Qwen3 32B Alibaba Cloud / Qwen Team 93.8 来源 ↗
3 Qwen3 30B A3B Alibaba Cloud / Qwen Team 91.0 来源 ↗
4 Llama-3.3 Nemotron Super 49B v1 NVIDIA 88.3 来源 ↗
5 Mistral Small 3 24B Instruct Mistral AI 87.6 来源 ↗
6 Qwen2.5 72B Instruct Alibaba Cloud / Qwen Team 81.2 来源 ↗
7 Phi 4 Reasoning Plus Microsoft 79.0 来源 ↗
8 DeepSeek-V2.5 DeepSeek 76.2 来源 ↗
9 Phi 4 Microsoft 75.4 来源 ↗
10 Phi 4 Reasoning Microsoft 73.3 来源 ↗
11 Ministral 8B Instruct Mistral AI 70.9 来源 ↗
12 Jamba 1.5 Large AI21 Labs 65.4 来源 ↗
13 Granite 3.3 8B Instruct IBM 57.6 来源 ↗
14 Granite 3.3 8B Base IBM 57.6 来源 ↗
15 Qwen2.5 7B Instruct Alibaba Cloud / Qwen Team 52.0 来源 ↗
16 Jamba 1.5 Mini AI21 Labs 46.1 来源 ↗
17 Mistral Small 3.2 24B Instruct Mistral AI 43.1 来源 ↗
18 Phi-3.5-MoE-instruct Microsoft 37.9 来源 ↗
19 Phi-3.5-mini-instruct Microsoft 37.0 来源 ↗
20 Phi 4 Mini Microsoft 32.8 来源 ↗
21 IBM Granite 4.0 Tiny Preview IBM 26.7 来源 ↗