EdgeBench - Optimization: leaderboard

Metric: Combinatorial optimization tasks: mean best-so-far score (0-100) after a 12-hour continuous agent run with evaluator feedback, three runs per task, averaged over the family's tasks; higher is better. Source: arxiv.org. Saturation forecast: Around January 2028. 5 models tracked.

Top models

#ModelScore
1Claude Opus 4.836.5
2GPT-5.533.6
3GPT-5.427.9
4GLM-5.126.4
5DeepSeek V4 Pro21.5

Interactive version: theaggregate.ai/benchmark?slug=edgebench-optimization · How It Works · Data refreshed daily, snapshot 2026-09-29.