WorldCoder-Bench - Application: leaderboard

Metric: Verification coverage (%) on the Application tasks (product, visualization, game, architecture, animation) of WorldCoder-Core, the 205-task hard evaluation split: from a natural-language task (optional .glb assets and a visible runtime-state interface) the model writes one self-contained Three.js HTML program zero-shot, and StateProbe runs it in a sandboxed browser and checks hidden, mutation-hardened behavioural contracts over runtime states and transitions; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScore
1GPT-5.423.5
2Qwen 3.6 Plus23.2
3Gemini 3.1 Pro (Preview)21.4
4MiniMax-M2.719.7
5DeepSeek V3.218.8
6Kimi K2.515.8
7DeepSeek V4 Flash15.8
8Claude Opus 4.614.6
9Claude Sonnet 4.613.7

Interactive version: theaggregate.ai/benchmark?slug=worldcoder-bench-application · How It Works · Data refreshed daily, snapshot 2026-09-29.