Benchmark

BIG-Bench Extra Hard

reasoning general language text

BIG-Bench Extra Hard (BBEH) is a challenging benchmark that replaces each task in BIG-Bench Hard with a novel task that probes similar reasoning capabilities but exhibits significantly increased difficulty. The benchmark contains 23 tasks testing diverse reasoning skills including many-hop reasoning, causal understanding, spatial reasoning, temporal arithmetic, geometric reasoning, linguistic reasoning, logic puzzles, and humor understanding. Designed to address saturation on existing benchmarks where state-of-the-art models achieve near-perfect scores, BBEH shows substantial room for improvement with best models achieving only 9.8-44.8% average accuracy.

语言EN
满分1
参评模型5

模型排名

名次 模型 机构 分数 来源
1 Gemma 3 27B Google 19.3 来源 ↗
2 Gemma 3 12B Google 16.3 来源 ↗
3 Gemini Diffusion Google 15.0 来源 ↗
4 Gemma 3 4B Google 11.0 来源 ↗
5 Gemma 3 1B Google 7.2 来源 ↗