Benchmark

BIG-Bench Hard

reasoning math language text

BIG-Bench Hard (BBH) is a subset of 23 challenging BIG-Bench tasks selected because prior language model evaluations did not outperform average human-rater performance. The benchmark contains 6,511 evaluation examples testing various forms of multi-step reasoning including arithmetic, logical reasoning (Boolean expressions, logical deduction), geometric reasoning, temporal reasoning, and language understanding. Tasks require capabilities such as causal judgment, object counting, navigation, pattern recognition, and complex problem solving.

语言EN
满分1
参评模型21

模型排名

名次 模型 机构 分数 来源
1 Claude 3.5 Sonnet Anthropic 93.1 来源 ↗
2 Claude 3.5 Sonnet Anthropic 93.1 来源 ↗
3 Gemini 1.5 Pro Google 89.2 来源 ↗
4 Gemma 3 27B Google 87.6 来源 ↗
5 Claude 3 Opus Anthropic 86.8 来源 ↗
6 Gemma 3 12B Google 85.7 来源 ↗
7 Gemini 1.5 Flash Google 85.5 来源 ↗
8 Claude 3 Sonnet Anthropic 82.9 来源 ↗
9 Phi-3.5-MoE-instruct Microsoft 79.1 来源 ↗
10 Claude 3 Haiku Anthropic 73.7 来源 ↗
11 Gemma 3 4B Google 72.2 来源 ↗
12 Phi 4 Mini Microsoft 70.4 来源 ↗
13 Granite 3.3 8B Instruct IBM 69.1 来源 ↗
14 Granite 3.3 8B Base IBM 69.1 来源 ↗
15 Phi-3.5-mini-instruct Microsoft 69.0 来源 ↗
16 IBM Granite 4.0 Tiny Preview IBM 55.7 来源 ↗
17 Gemma 3n E4B Instructed LiteRT Preview Google 52.9 来源 ↗
18 Gemma 3n E4B Google 52.9 来源 ↗
19 Gemma 3n E2B Instructed LiteRT (Preview) Google 44.3 来源 ↗
20 Gemma 3n E2B Google 44.3 来源 ↗
21 Gemma 3 1B Google 39.1 来源 ↗