WolfBench (Codex): leaderboard
Metric: Mean Success Rate (%, 89 Terminal-Bench 2.0 tasks). Source: wolfbench.ai. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (WolfBench Run: agent=codex; version=0.144.4; timeout=3600s; thinking=max; providers=openai; source-model=GPT-5.6%20Sol) | 86.74 |
| 2 | gpt-6-astra (WolfBench Run: agent=codex; version=0.153.3; timeout=3600s; thinking=max; providers=openai; source-model=gpt-6-astra) | 85.84 |
| 3 | GPT-5.5 (WolfBench Run: agent=codex; version=0.125.0; timeout=3600s; thinking=high; providers=openai; source-model=GPT-5.5) | 78.65 |
Interactive version: theaggregate.ai/benchmark?slug=wolfbench-codex · How It Works · Data refreshed daily, snapshot 2026-10-09.