EnvSimBench - Simple: leaderboard

Metric: Config Match (%): share of the 50 Simple state-changing samples (1 or 2 changed state fields) whose predicted state changes, applied to the before-state, reproduce the executor after-state exactly; MDP prompt with the state, the action and its code, non-thinking mode; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.

Top models

#ModelScore
1MiniMax-M2.750
2Gemini 3.1 Pro (Preview)48
3Qwen 3.5 397B A17B (Non-reasoning)46
4GLM-5 (Non-reasoning)44
5Qwen 3.5 9B (Non-reasoning)42
6Qwen 3.5 4B (Non-reasoning)42
7Claude Sonnet 4.640
8GPT-5.4 (Non-reasoning)40
9Qwen 3.5 27B (Non-reasoning)26
10DeepSeek V3.2 (Non-reasoning)22

Interactive version: theaggregate.ai/benchmark?slug=envsimbench-simple · How It Works · Data refreshed daily, snapshot 2026-10-07.