Benchmark

BIG-Bench

reasoning math language text 多语言

Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark consisting of 204+ tasks designed to probe large language models and extrapolate their future capabilities. It covers diverse domains including linguistics, mathematics, common-sense reasoning, biology, physics, social bias, software development, and more. The benchmark focuses on tasks believed to be beyond current language model capabilities and includes both English and non-English tasks across multiple languages.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Gemini 1.0 Pro Google 75.0 来源 ↗
2 Gemma 2 27B Google 74.9 来源 ↗
3 Gemma 2 9B Google 68.2 来源 ↗