WolfBench (OpenClaw): leaderboard
Metric: Mean Success Rate (%, 89 Terminal-Bench 2.0 tasks). Source: wolfbench.ai. Saturation forecast: Rough model projection: around 2026. 26 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 (WolfBench Run: agent=openclaw; version=2026.3.11; timeout=3600s; thinking=off; providers=anthropic; source-model=Claude%20Opus%204.7) | 75 |
| 2 | Gemini 3.5 Flash (WolfBench Run: agent=openclaw; version=2026.4.23; timeout=3600s; thinking=high; providers=google; source-model=Gemini%203.5%20Flash) | 72.19 |
| 3 | GPT-5.4 (WolfBench Run: agent=openclaw; version=2026.3.11; timeout=3600s; thinking=xhigh; providers=openai; source-model=GPT-5.4) | 71.01 |
| 4 | GPT-5.5 (WolfBench Run: agent=openclaw; version=2026.4.23; timeout=3600s; thinking=off; providers=openai; source-model=GPT-5.5) | 69.38 |
| 5 | Gemini 3.5 Flash (WolfBench Run: agent=openclaw; version=2026.4.23; timeout=3600s; thinking=medium; providers=google; source-model=Gemini%203.5%20Flash) | 69.38 |
| 6 | Gemini 3.5 Flash (WolfBench Run: agent=openclaw; version=2026.4.23; timeout=3600s; thinking=low; providers=google; source-model=Gemini%203.5%20Flash) | 67.42 |
| 7 | Claude Sonnet 5 (WolfBench Run: agent=openclaw; version=2026.3.1; timeout=3600s; thinking=high; providers=anthropic; source-model=Claude%20Sonnet%205) | 61.8 |
| 8 | GPT-5.4 (WolfBench Run: agent=openclaw; version=2026.3.11; timeout=3600s; thinking=low; providers=openai; source-model=GPT-5.4) | 61.24 |
| 9 | Claude Opus 4.6 (WolfBench Run: agent=openclaw; version=2026.3.11; timeout=3600s; thinking=max; providers=anthropic; source-model=Claude%20Opus%204.6) | 58.99 |
| 10 | Kimi K2.6 [Moonshot AI] (WolfBench Run: agent=openclaw; version=2026.3.11; timeout=3600s; thinking=-; providers=openrouter; source-model=Kimi%20K2.6%20%5BMoonshot%20AI%5D) | 58.43 |
Interactive version: theaggregate.ai/benchmark?slug=wolfbench-openclaw · How It Works · Data refreshed daily, snapshot 2026-10-09.