AlgoWorlds - Reference Utility: leaderboard
Metric: Reference utility (%; x100 of a 0-1 score per feasible decision, 1 at the global optimum and 0 at a predefined suboptimal reference decision, 0 when infeasible; 240 partially observed environments built from ten combinatorial optimization families at four workload levels; the agent sees a hidden instance only through task-specific information tools, then commits one structured decision checked by an independent verifier; every model at its maximum reasoning setting; mean of three trials). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (Max) | 74.98 |
| 2 | Claude Opus 4.8 (Max) | 73.14 |
| 3 | Claude Sonnet 5 (Max) | 67.2 |
| 4 | GPT-5.6 Terra (Max) | 54.1 |
| 5 | GLM-5.2 (Max) | 35.48 |
| 6 | Qwen 3.5 Plus | 15.82 |
Interactive version: theaggregate.ai/benchmark?slug=algoworlds-reference-utility · How It Works · Data refreshed daily, snapshot 2026-09-26.