Benchmark
Social IQa
reasoning
psychology
text
The first large-scale benchmark for commonsense reasoning about social situations. Contains 38,000 multiple choice questions probing emotional and social intelligence in everyday situations, testing commonsense understanding of social interactions and theory of mind reasoning about the implied emotions and behavior of others.
语言EN
满分1
参评模型9
模型排名
| 名次 | 模型 | 机构 | 分数 | 来源 |
|---|---|---|---|---|
| 1 | Phi-3.5-MoE-instruct | Microsoft | 78.0 | 来源 ↗ |
| 2 | Phi-3.5-mini-instruct | Microsoft | 74.7 | 来源 ↗ |
| 3 | Phi 4 Mini | Microsoft | 72.5 | 来源 ↗ |
| 4 | Gemma 2 27B | 53.7 | 来源 ↗ | |
| 5 | Gemma 2 9B | 53.4 | 来源 ↗ | |
| 6 | Gemma 3n E4B | 50.0 | 来源 ↗ | |
| 7 | Gemma 3n E4B Instructed LiteRT Preview | 50.0 | 来源 ↗ | |
| 8 | Gemma 3n E2B | 48.8 | 来源 ↗ | |
| 9 | Gemma 3n E2B Instructed LiteRT (Preview) | 48.8 | 来源 ↗ |