Benchmark

Winogrande

reasoning language text

WinoGrande: An Adversarial Winograd Schema Challenge at Scale. A large-scale dataset of 44,000 pronoun resolution problems designed to test machine commonsense reasoning. Uses adversarial filtering to reduce spurious biases and provides a more robust evaluation of whether AI systems truly understand commonsense or exploit statistical shortcuts. Current best AI methods achieve 59.4-79.1% accuracy, significantly below human performance of 94.0%.

语言EN
满分1
参评模型19

模型排名

名次 模型 机构 分数 来源
1 GPT-4 OpenAI 87.5 来源 ↗
2 Command R+ Cohere 85.4 来源 ↗
3 Qwen2 72B Instruct Alibaba Cloud / Qwen Team 85.1 来源 ↗
4 Llama 3.1 Nemotron 70B Instruct NVIDIA 84.5 来源 ↗
5 Gemma 2 27B Google 83.7 来源 ↗
6 Qwen2.5 32B Instruct Alibaba Cloud / Qwen Team 82.0 来源 ↗
7 Phi-3.5-MoE-instruct Microsoft 81.3 来源 ↗
8 Qwen2.5-Coder 32B Instruct Alibaba Cloud / Qwen Team 80.8 来源 ↗
9 Gemma 2 9B Google 80.6 来源 ↗
10 Mistral NeMo Instruct Mistral AI 76.8 来源 ↗
11 Ministral 8B Instruct Mistral AI 75.3 来源 ↗
12 Granite 3.3 8B Base IBM 74.4 来源 ↗
13 Qwen2.5-Coder 7B Instruct Alibaba Cloud / Qwen Team 72.9 来源 ↗
14 Gemma 3n E4B Instructed LiteRT Preview Google 71.7 来源 ↗
15 Gemma 3n E4B Google 71.7 来源 ↗
16 Phi-3.5-mini-instruct Microsoft 68.5 来源 ↗
17 Phi 4 Mini Microsoft 67.0 来源 ↗
18 Gemma 3n E2B Google 66.8 来源 ↗
19 Gemma 3n E2B Instructed LiteRT (Preview) Google 66.8 来源 ↗