EnvSimBench - Medium: leaderboard

Metric: Config Match (%): share of the 200 Medium state-changing samples (3 to 6 changed state fields) whose predicted state changes, applied to the before-state, reproduce the executor after-state exactly; MDP prompt with the state, the action and its code, non-thinking mode; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)17.5
2GPT-5.4 (Non-reasoning)17.5
3Qwen 3.5 397B A17B (Non-reasoning)17
4MiniMax-M2.716
5GLM-5 (Non-reasoning)14
6Claude Sonnet 4.612
7Qwen 3.5 9B (Non-reasoning)11.6
8Qwen 3.5 4B (Non-reasoning)10.5
9DeepSeek V3.2 (Non-reasoning)8.5
10Qwen 3.5 27B (Non-reasoning)2

Interactive version: theaggregate.ai/benchmark?slug=envsimbench-medium · How It Works · Data refreshed daily, snapshot 2026-10-07.