PIQA: leaderboard

Physical Intuition QA: 20K multiple-choice questions testing physical commonsense reasoning - understanding how everyday objects and actions work in the physical world.

Metric: Accuracy (%). Source: epoch.ai. Status: saturated. 60 models tracked.

Top models

#ModelScore
1Median Human95
2GPT-4o Mini (2024-07-18)88.7
3Phi-3.5-MoE-instruct88.6
4Gemini 1.5 Flash (002)87.5
5Llama 3.1 405B85.9
6DeepSeek V283.9
7Gemma 2 9B83.7
8Mixtral 8x7B (v0.1)83.6
9Mistral-Nemo-Base-240783.5
10Mistral-7B-v0.183
11falcon-40B83
12Llama 2 70B Base82.8
13LLaMA-65B82.8
14Qwen 2.5 72B82.6
15text-davinci-00182.3

Interactive version: theaggregate.ai/benchmark?slug=piqa · How It Works · Data refreshed daily, snapshot 2026-09-05.