Benchmark

ZebraLogic

reasoning text

ZebraLogic is an evaluation framework for assessing large language models' logical reasoning capabilities through logic grid puzzles derived from constraint satisfaction problems (CSPs). The benchmark consists of 1,000 programmatically generated puzzles with controllable and quantifiable complexity, revealing a 'curse of complexity' where model accuracy declines significantly as problem complexity grows.

语言EN
满分1
参评模型3

模型排名

名次 模型 机构 分数 来源
1 Qwen3-235B-A22B-Instruct-2507 Alibaba Cloud / Qwen Team 95.0 来源 ↗
2 Kimi K2 Instruct Moonshot AI 89.0 来源 ↗
3 Kimi K2-Instruct-0905 Moonshot AI 89.0 来源 ↗