EnvSimBench - Difficult: leaderboard

Metric: Config Match (%): share of the 50 Difficult state-changing samples (7 to 12 changed state fields) whose predicted state changes, applied to the before-state, reproduce the executor after-state exactly; MDP prompt with the state, the action and its code, non-thinking mode; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.

Top models

#ModelScore
1GLM-5 (Non-reasoning)28
2GPT-5.4 (Non-reasoning)26
3Qwen 3.5 397B A17B (Non-reasoning)24
4MiniMax-M2.722
5Qwen 3.5 9B (Non-reasoning)20
6Gemini 3.1 Pro (Preview)18
7Claude Sonnet 4.616
8Qwen 3.5 4B (Non-reasoning)6
9DeepSeek V3.2 (Non-reasoning)4
10Qwen 3.5 27B (Non-reasoning)4

Interactive version: theaggregate.ai/benchmark?slug=envsimbench-difficult · How It Works · Data refreshed daily, snapshot 2026-10-07.