Benchmark

PIQA

reasoning physics general text

PIQA (Physical Interaction: Question Answering) is a benchmark dataset for physical commonsense reasoning in natural language. It tests AI systems' ability to answer questions requiring physical world knowledge through multiple choice questions with everyday situations, focusing on atypical solutions inspired by instructables.com. The dataset contains 21,000 multiple choice questions where models must choose the most appropriate solution for physical interactions.

语言EN
满分1
参评模型9

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-MoE-instruct Microsoft 88.6 来源 ↗
2 Gemma 2 27B Google 83.2 来源 ↗
3 Gemma 2 9B Google 81.7 来源 ↗
4 Gemma 3n E4B Google 81.0 来源 ↗
5 Gemma 3n E4B Instructed LiteRT Preview Google 81.0 来源 ↗
6 Phi-3.5-mini-instruct Microsoft 81.0 来源 ↗
7 Gemma 3n E2B Google 78.9 来源 ↗
8 Gemma 3n E2B Instructed LiteRT (Preview) Google 78.9 来源 ↗
9 Phi 4 Mini Microsoft 77.6 来源 ↗