PIQA — leaderboard

Physical Intuition QA: 20K multiple-choice questions testing physical commonsense reasoning - understanding how everyday objects and actions work in the physical world.

Metric: Accuracy (%). Source: epoch.ai. Status: saturation imminent. 60 models tracked.

Top models

#ModelScore
1Median Human95
2GPT-4o Mini (2024-07-18)88.7
3Phi-3.5-MoE-instruct88.6
4Gemini 1.5 Flash (002)87.5
5Llama 3.1 405B85.9
6falcon-180B84.9
7DeepSeek V283.9
8Gemma 2 9B83.7
9Mixtral 8x7B (v0.1)83.6
10Mistral-Nemo-Base-240783.5
11StableBeluga283.3
12Mistral-7B-v0.183
13falcon-40B83
14Llama 2 70B Base82.8
15LLaMA-65B82.8

Interactive version: theaggregate.ai/benchmark?slug=piqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.