CausalPhys: leaderboard

Metric: Answer accuracy (0-1 scaled to %) over all four domains (Perception, Anticipation, Intervention, Goal-Orientation; 3,062 image- and video-based questions with verified answers over real-world physical scenes), item-weighted; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 11 models tracked.

Top models

#ModelScore
1InternVL3-78B59.24
2GPT-4o58.88
3Gemini 2.5 Flash57.25
4GPT-4o Mini56.3
5Claude Sonnet 454.28
6Qwen 3 VL 32B52.68
7Phi-4 Multimodal Instruct51.99
8Qwen 2 VL 7B50.65
9Mistral Small 3.246.11

Interactive version: theaggregate.ai/benchmark?slug=causalphys · How It Works · Data refreshed daily, snapshot 2026-09-29.