CursorBench 4.0: leaderboard

Cursor coding-agent benchmark for ambiguous, multi-file real-world tasks; the 4.0 task set adds long-horizon editing, refactoring, investigation, intent understanding, job management, and design adherence.

Metric: Score (%). Source: cursor.com. Status: years away from saturation. 39 models tracked.

Top models

#ModelScore
1Claude Fable 5.1 (Max)51.8
2Claude Fable 5.1 (xHigh)51.6
3Claude Fable 5.1 (High)49.2
4Claude Fable 5.1 (Medium)46.8
5Claude Opus 546.6
6Claude Fable 5.1 (Low)45.1
7GPT-5.6 Sol (Max)41.7
8Muse Spark 1.3 (Max)41.6
9Grok 4.6 (xHigh)41.4
10GPT-5.6 Terra (Max)41.3
11Grok 4.6 (High)40.4
12Gemini 3.8 Flash (High)39.6
13GPT-5.6 Sol (xHigh)37.7
14Gemini 3.8 Flash (Medium)37.3
15Grok 4.6 (Medium)36.1

Interactive version: theaggregate.ai/benchmark?slug=cursorbench-4-0 · How It Works · Data refreshed daily, snapshot 2026-09-19.