WolfBench (Cursor CLI): leaderboard
Metric: Mean Success Rate (%, 89 Terminal-Bench 2.0 tasks). Source: wolfbench.ai. 1 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 (WolfBench Run: agent=cursor-cli; version=2026.04.17; timeout=3600s; thinking=high; providers=cursor; source-model=GPT-5.5) | 77.9 |
Interactive version: theaggregate.ai/benchmark?slug=wolfbench-cursor-cli · How It Works · Data refreshed daily, snapshot 2026-10-09.