SimuScene: leaderboard

Metric: Avg@8 accuracy (%): mean share of responses whose animation passes every verification question; the model writes Python code for an animation of each of SimuScene's 334 human-verified physics scenarios (52 concepts in mechanics, electromagnetism, optics, fluid mechanics and thermodynamics); the code is run, the video rendered, and Qwen2.5-VL-72B-Instruct answers the scenario's visual verification questions; a response counts only when every question is answered true; eight responses per scenario, default settings with thinking enabled; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek R1 052821.5#217
2GPT-5 (Medium)20.5#91 (GPT-5)
3O4 Mini17.2#172
4O315.9#121
5Qwen 3 235B A22B (Thinking)15.1#304 (Qwen 3 235B A22B)
6DeepSeek V3.1 (Thinking)14.5#260 (DeepSeek V3.1)
7GPT-OSS-120B14#330
8Gemini 2.5 Pro12.7#145
9Qwen 3 32B (Thinking)11.1#424 (Qwen 3 32B)
10GPT-OSS-20B10.5#499

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=simuscene · How It Works · Data refreshed daily, snapshot 2026-10-11.