Benchmark

BBH

reasoning math language text

Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average human-rater. These tasks require multi-step reasoning across diverse domains including arithmetic, logical reasoning, reading comprehension, and commonsense reasoning. The benchmark was designed to test capabilities believed to be beyond current language models and focuses on evaluating complex reasoning skills including temporal understanding, spatial reasoning, causal understanding, and deductive logical reasoning.

语言EN
满分1
参评模型8

模型排名

名次 模型 机构 分数 来源
1 Qwen3 235B A22B Alibaba Cloud / Qwen Team 88.9 来源 ↗
2 Nova Pro Amazon 86.9 来源 ↗
3 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 84.5 来源 ↗
4 DeepSeek-V2.5 DeepSeek 84.3 来源 ↗
5 Nova Lite Amazon 82.4 来源 ↗
6 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 82.4 来源 ↗
7 Nova Micro Amazon 79.5 来源 ↗
8 Qwen2.5 14B Instruct Alibaba Cloud / Qwen Team 78.2 来源 ↗