Benchmark

LiveBench 20241125

math reasoning general text

LiveBench is a challenging, contamination-limited LLM benchmark that addresses test set contamination by releasing new questions monthly based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. It comprises tasks across math, coding, reasoning, language, instruction following, and data analysis with verifiable, objective ground-truth answers.

语言EN
满分1
参评模型4

模型排名

名次 模型 机构 分数 来源
1 Qwen3-235B-A22B-Thinking-2507 Alibaba Cloud / Qwen Team 78.4 来源 ↗
2 Qwen3-Next-80B-A3B-Thinking Alibaba Cloud / Qwen Team 76.6 来源 ↗
3 Qwen3-Next-80B-A3B-Instruct Alibaba Cloud / Qwen Team 75.8 来源 ↗
4 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 75.4 来源 ↗