Benchmark

Arena-Hard v2

general reasoning creativity text

Arena-Hard-Auto v2 is a challenging benchmark consisting of 500 carefully curated prompts sourced from Chatbot Arena and WildChat-1M, designed to evaluate large language models on real-world user queries. The benchmark covers diverse domains including open-ended software engineering problems, mathematics, creative writing, and technical problem-solving. It uses LLM-as-a-Judge for automatic evaluation, achieving 98.6% correlation with human preference rankings while providing 3x higher separation of model performances compared to MT-Bench. The benchmark emphasizes prompt specificity, complexity, and domain knowledge to better distinguish between model capabilities.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen3-Next-80B-A3B-Instruct Alibaba Cloud / Qwen Team 82.7 来源 ↗
2 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 79.7 来源 ↗
3 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 79.2 来源 ↗
4 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 62.3 来源 ↗