WorldCoder-Bench - Simulation: leaderboard

Metric: Verification coverage (%) on the Simulation tasks (physics, chemistry, soft-body, molecular and complex-system worlds) of WorldCoder-Core, the 205-task hard evaluation split: from a natural-language task (optional .glb assets and a visible runtime-state interface) the model writes one self-contained Three.js HTML program zero-shot, and StateProbe runs it in a sandboxed browser and checks hidden, mutation-hardened behavioural contracts over runtime states and transitions; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)27
2GPT-5.426.8
3Qwen 3.6 Plus24.6
4MiniMax-M2.722.6
5Claude Sonnet 4.621.4
6DeepSeek V3.221.3
7Claude Opus 4.618.2
8Kimi K2.513.6
9DeepSeek V4 Flash11.8

Interactive version: theaggregate.ai/benchmark?slug=worldcoder-bench-simulation · How It Works · Data refreshed daily, snapshot 2026-09-29.