WorldCoder-Bench: leaderboard

Metric: Verification coverage (%; share of hidden assertions that pass per task, averaged over tasks) on WorldCoder-Core, the 205-task hard evaluation split: from a natural-language task (optional .glb assets and a visible runtime-state interface) the model writes one self-contained Three.js HTML program zero-shot, and StateProbe runs it in a sandboxed browser and checks hidden, mutation-hardened behavioural contracts over runtime states and transitions; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1GPT-5.427.8
2Gemini 3.1 Pro (Preview)26.5
3Qwen 3.6 Plus25.3
4MiniMax-M2.723
5DeepSeek V3.221.9
6Claude Sonnet 4.618.7
7Kimi K2.518.2
8Claude Opus 4.617.5
9DeepSeek V4 Flash16.5

Interactive version: theaggregate.ai/benchmark?slug=worldcoder-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.