PhysCodeBench: leaderboard

Metric: PhysCodeEval total score (0-100): code quality (50 points for executing and writing the expected simulation file) plus simulation fidelity (50 points for CLIPScore alignment of the rendered video with the instruction and flow-based motion smoothness) of Genesis physics-simulation programs generated zero-shot from natural-language scene descriptions on the 200-example PhysCodeBench test set (mean of five attempts, temperature 0.1, 4,096-token cap), with the engine documentation in context (the full 100K-token corpus for long-context models, 20K tokens of BM25-retrieved sections for 32K-context models); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet35.9
2GPT-4o33.8
3DeepSeek R129.5
4DeepSeek R1 Distill Qwen 32B27.6
5QwQ-32B15.5

Interactive version: theaggregate.ai/benchmark?slug=physcodebench · How It Works · Data refreshed daily, snapshot 2026-10-07.