WinoGrande — leaderboard

Large-scale Winograd Schema Challenge with 44K problems testing commonsense reasoning through pronoun resolution. Models must determine what a pronoun refers to using world knowledge.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 80 models tracked.

Top models

#ModelScore
1Median Human94
2Llama 3.1 405B89.2
3Claude 3 Opus (20240229)88.5
4GPT-4 (0314)87.5
5falcon-180B87.1
6DeepSeek V286.3
7DeepSeek V385.2
8Llama 3 70B83.5
9Qwen 2.5 72B82.3
10Llama 3.1 405B Instruct82.2
11GPT-3.5 Turbo (0613)81.6
12text-davinci-00281.6
13Phi-3-small-8k-instruct81.5
14Phi-3-medium-128k-instruct81.5
15Qwen 2.5 Coder 32B80.8

Interactive version: theaggregate.ai/benchmark?slug=winogrande · How the rankings work · Data refreshed daily, snapshot 2026-07-22.