AlgoWorlds - Feasibility: leaderboard
Metric: Feasibility (%; share of worlds where the committed decision satisfies every task constraint; 240 partially observed environments built from ten combinatorial optimization families at four workload levels; the agent sees a hidden instance only through task-specific information tools, then commits one structured decision checked by an independent verifier; every model at its maximum reasoning setting; mean of three trials). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (Max) | 97.5 |
| 2 | Claude Opus 4.8 (Max) | 96.25 |
| 3 | Claude Sonnet 5 (Max) | 92.08 |
| 4 | GPT-5.6 Terra (Max) | 87.78 |
| 5 | GLM-5.2 (Max) | 64.58 |
| 6 | Qwen 3.5 Plus | 48.89 |
Interactive version: theaggregate.ai/benchmark?slug=algoworlds-feasibility · How It Works · Data refreshed daily, snapshot 2026-09-26.