Benchmark

Social IQa

reasoning psychology text

The first large-scale benchmark for commonsense reasoning about social situations. Contains 38,000 multiple choice questions probing emotional and social intelligence in everyday situations, testing commonsense understanding of social interactions and theory of mind reasoning about the implied emotions and behavior of others.

语言EN
满分1
参评模型9

模型排名

名次 模型 机构 分数 来源
1 Phi-3.5-MoE-instruct Microsoft 78.0 来源 ↗
2 Phi-3.5-mini-instruct Microsoft 74.7 来源 ↗
3 Phi 4 Mini Microsoft 72.5 来源 ↗
4 Gemma 2 27B Google 53.7 来源 ↗
5 Gemma 2 9B Google 53.4 来源 ↗
6 Gemma 3n E4B Google 50.0 来源 ↗
7 Gemma 3n E4B Instructed LiteRT Preview Google 50.0 来源 ↗
8 Gemma 3n E2B Google 48.8 来源 ↗
9 Gemma 3n E2B Instructed LiteRT (Preview) Google 48.8 来源 ↗