PhysCodeBench: leaderboard
Metric: PhysCodeEval total score (0-100): code quality (50 points for executing and writing the expected simulation file) plus simulation fidelity (50 points for CLIPScore alignment of the rendered video with the instruction and flow-based motion smoothness) of Genesis physics-simulation programs generated zero-shot from natural-language scene descriptions on the 200-example PhysCodeBench test set (mean of five attempts, temperature 0.1, 4,096-token cap), with the engine documentation in context (the full 100K-token corpus for long-context models, 20K tokens of BM25-retrieved sections for 32K-context models); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude 3.5 Sonnet | 35.9 |
| 2 | GPT-4o | 33.8 |
| 3 | DeepSeek R1 | 29.5 |
| 4 | DeepSeek R1 Distill Qwen 32B | 27.6 |
| 5 | QwQ-32B | 15.5 |
Interactive version: theaggregate.ai/benchmark?slug=physcodebench · How It Works · Data refreshed daily, snapshot 2026-10-07.