Benchmark
BIG-Bench Hard
reasoning
math
language
text
BIG-Bench Hard (BBH) is a subset of 23 challenging BIG-Bench tasks selected because prior language model evaluations did not outperform average human-rater performance. The benchmark contains 6,511 evaluation examples testing various forms of multi-step reasoning including arithmetic, logical reasoning (Boolean expressions, logical deduction), geometric reasoning, temporal reasoning, and language understanding. Tasks require capabilities such as causal judgment, object counting, navigation, pattern recognition, and complex problem solving.
语言EN
满分1
参评模型21
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Claude 3.5 Sonnet | Anthropic | 93.1 | 来源 ↗ |
| 2 | Claude 3.5 Sonnet | Anthropic | 93.1 | 来源 ↗ |
| 3 | Gemini 1.5 Pro | 89.2 | 来源 ↗ | |
| 4 | Gemma 3 27B | 87.6 | 来源 ↗ | |
| 5 | Claude 3 Opus | Anthropic | 86.8 | 来源 ↗ |
| 6 | Gemma 3 12B | 85.7 | 来源 ↗ | |
| 7 | Gemini 1.5 Flash | 85.5 | 来源 ↗ | |
| 8 | Claude 3 Sonnet | Anthropic | 82.9 | 来源 ↗ |
| 9 | Phi-3.5-MoE-instruct | Microsoft | 79.1 | 来源 ↗ |
| 10 | Claude 3 Haiku | Anthropic | 73.7 | 来源 ↗ |
| 11 | Gemma 3 4B | 72.2 | 来源 ↗ | |
| 12 | Phi 4 Mini | Microsoft | 70.4 | 来源 ↗ |
| 13 | Granite 3.3 8B Instruct | IBM | 69.1 | 来源 ↗ |
| 14 | Granite 3.3 8B Base | IBM | 69.1 | 来源 ↗ |
| 15 | Phi-3.5-mini-instruct | Microsoft | 69.0 | 来源 ↗ |
| 16 | IBM Granite 4.0 Tiny Preview | IBM | 55.7 | 来源 ↗ |
| 17 | Gemma 3n E4B Instructed LiteRT Preview | 52.9 | 来源 ↗ | |
| 18 | Gemma 3n E4B | 52.9 | 来源 ↗ | |
| 19 | Gemma 3n E2B Instructed LiteRT (Preview) | 44.3 | 来源 ↗ | |
| 20 | Gemma 3n E2B | 44.3 | 来源 ↗ | |
| 21 | Gemma 3 1B | 39.1 | 来源 ↗ |