CursorBench 4.0: leaderboard
Cursor coding-agent benchmark for ambiguous, multi-file real-world tasks; the 4.0 task set adds long-horizon editing, refactoring, investigation, intent understanding, job management, and design adherence.
Metric: Score (%). Source: cursor.com. Status: years away from saturation. 39 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5.1 (Max) | 51.8 |
| 2 | Claude Fable 5.1 (xHigh) | 51.6 |
| 3 | Claude Fable 5.1 (High) | 49.2 |
| 4 | Claude Fable 5.1 (Medium) | 46.8 |
| 5 | Claude Opus 5 | 46.6 |
| 6 | Claude Fable 5.1 (Low) | 45.1 |
| 7 | GPT-5.6 Sol (Max) | 41.7 |
| 8 | Muse Spark 1.3 (Max) | 41.6 |
| 9 | Grok 4.6 (xHigh) | 41.4 |
| 10 | GPT-5.6 Terra (Max) | 41.3 |
| 11 | Grok 4.6 (High) | 40.4 |
| 12 | Gemini 3.8 Flash (High) | 39.6 |
| 13 | GPT-5.6 Sol (xHigh) | 37.7 |
| 14 | Gemini 3.8 Flash (Medium) | 37.3 |
| 15 | Grok 4.6 (Medium) | 36.1 |
Interactive version: theaggregate.ai/benchmark?slug=cursorbench-4-0 · How It Works · Data refreshed daily, snapshot 2026-09-19.