Benchmark
BIG-Bench Extra Hard
reasoning
general
language
text
BIG-Bench Extra Hard (BBEH) is a challenging benchmark that replaces each task in BIG-Bench Hard with a novel task that probes similar reasoning capabilities but exhibits significantly increased difficulty. The benchmark contains 23 tasks testing diverse reasoning skills including many-hop reasoning, causal understanding, spatial reasoning, temporal arithmetic, geometric reasoning, linguistic reasoning, logic puzzles, and humor understanding. Designed to address saturation on existing benchmarks where state-of-the-art models achieve near-perfect scores, BBEH shows substantial room for improvement with best models achieving only 9.8-44.8% average accuracy.
语言EN
满分1
参评模型5
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Gemma 3 27B | 19.3 | 来源 ↗ | |
| 2 | Gemma 3 12B | 16.3 | 来源 ↗ | |
| 3 | Gemini Diffusion | 15.0 | 来源 ↗ | |
| 4 | Gemma 3 4B | 11.0 | 来源 ↗ | |
| 5 | Gemma 3 1B | 7.2 | 来源 ↗ |