Benchmark

AutoLogi

reasoning text 多语言

AutoLogi is an automated method for synthesizing open-ended logic puzzles to evaluate reasoning abilities of Large Language Models. The benchmark addresses limitations of existing multiple-choice reasoning evaluations by featuring program-based verification and controllable difficulty levels. It includes 1,575 English and 883 Chinese puzzles, enabling more reliable evaluation that better distinguishes models' reasoning capabilities across languages.

语言EN
满分1
参评模型2

模型排名

名次 模型 机构 分数 来源
1 Kimi K2 Instruct Moonshot AI 89.5 来源 ↗
2 Kimi K2-Instruct-0905 Moonshot AI 89.5 来源 ↗