Evo-Bench - Office: leaderboard

Metric: Held-out score (%; the fixed DeepSeek-V4-Flash policy runs the harness the model evolved from a CodeAct seed within 20 iterations, 1,000 steps and 48 hours; mean of 64 GDPval and 64 APEX-Agents held-out tasks; one run). Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (Max)41.6
2Gemma 4 31B (IT) (Thinking)40.4
3Claude Opus 4.8 (Max)39.7
4GLM-5.2 (Max)39.2
5DeepSeek V4 Pro (Max)39.1

Interactive version: theaggregate.ai/benchmark?slug=evo-bench-office · How It Works · Data refreshed daily, snapshot 2026-09-26.