WolfBench (Claude Code): leaderboard
Metric: Mean Success Rate (%, 89 Terminal-Bench 2.0 tasks). Source: wolfbench.ai. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 5 (WolfBench Run: agent=claude-code; version=2.1.197; timeout=3600s; thinking=-; providers=anthropic; source-model=Claude%20Sonnet%205) | 74.16 |
| 2 | Claude Opus 4.7 (WolfBench Run: agent=claude-code; version=2.1.112; timeout=3600s; thinking=xhigh; providers=anthropic; source-model=Claude%20Opus%204.7) | 73.31 |
| 3 | Claude Opus 4.6 (WolfBench Run: agent=claude-code; version=2.1.63; timeout=3600s; thinking=high; providers=anthropic; source-model=Claude%20Opus%204.6) | 63.3 |
| 4 | Claude Sonnet 4.6 (WolfBench Run: agent=claude-code; version=2.1.63; timeout=3600s; thinking=-; providers=anthropic; source-model=Claude%20Sonnet%204.6) | 57.87 |
Interactive version: theaggregate.ai/benchmark?slug=wolfbench-claude-code · How It Works · Data refreshed daily, snapshot 2026-10-09.