Benchmark
BBH
reasoning
math
language
text
Big-Bench Hard (BBH) is a suite of 23 challenging tasks selected from BIG-Bench for which prior language model evaluations did not outperform the average human-rater. These tasks require multi-step reasoning across diverse domains including arithmetic, logical reasoning, reading comprehension, and commonsense reasoning. The benchmark was designed to test capabilities believed to be beyond current language models and focuses on evaluating complex reasoning skills including temporal understanding, spatial reasoning, causal understanding, and deductive logical reasoning.
语言EN
满分1
参评模型8
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Qwen3 235B A22B | Alibaba Cloud / Qwen Team | 88.9 | 来源 ↗ |
| 2 | Nova Pro | Amazon | 86.9 | 来源 ↗ |
| 3 | Qwen2.5 32B Instruct | Alibaba Cloud / Qwen Team | 84.5 | 来源 ↗ |
| 4 | DeepSeek-V2.5 | DeepSeek | 84.3 | 来源 ↗ |
| 5 | Nova Lite | Amazon | 82.4 | 来源 ↗ |
| 6 | Qwen2 72B Instruct | Alibaba Cloud / Qwen Team | 82.4 | 来源 ↗ |
| 7 | Nova Micro | Amazon | 79.5 | 来源 ↗ |
| 8 | Qwen2.5 14B Instruct | Alibaba Cloud / Qwen Team | 78.2 | 来源 ↗ |