WorldCoder-Bench - Rendering: leaderboard

Metric: Verification coverage (%) on the Rendering tasks (graphics, materials, creative, shaders, post-processing) of WorldCoder-Core, the 205-task hard evaluation split: from a natural-language task (optional .glb assets and a visible runtime-state interface) the model writes one self-contained Three.js HTML program zero-shot, and StateProbe runs it in a sandboxed browser and checks hidden, mutation-hardened behavioural contracts over runtime states and transitions; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1GPT-5.437.8
2Gemini 3.1 Pro (Preview)35.9
3Qwen 3.6 Plus30.4
4MiniMax-M2.730.2
5Kimi K2.529.3
6DeepSeek V3.228.8
7Claude Sonnet 4.625.1
8DeepSeek V4 Flash24.2
9Claude Opus 4.622.2

Interactive version: theaggregate.ai/benchmark?slug=worldcoder-bench-rendering · How It Works · Data refreshed daily, snapshot 2026-09-29.