WinoGrande: leaderboard

Large-scale Winograd Schema Challenge with 44K problems testing commonsense reasoning through pronoun resolution. Models must determine what a pronoun refers to using world knowledge.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 80 models tracked.

Top models

#ModelScore
1Median Human94
2Llama 3.1 405B89.2
3Claude 3 Opus (20240229)88.5
4GPT-4 (0314)87.5
5DeepSeek V286.3
6DeepSeek V385.2
7Llama 3 70B83.5
8Qwen 2.5 72B82.3
9Llama 3.1 405B Instruct82.2
10GPT-3.5 Turbo (0613)81.6
11text-davinci-00281.6
12Phi-3-small-8k-instruct81.5
13Phi-3-medium-128k-instruct81.5
14Qwen 2.5 Coder 32B80.8
15Llama 2 70B Base80.2

Interactive version: theaggregate.ai/benchmark?slug=winogrande · How It Works · Data refreshed daily, snapshot 2026-09-05.