FrontierCode — leaderboard

Cognition benchmark for production-quality coding agents measuring whether maintainers would merge model PRs. It uses 150 maintainer-authored open source tasks with nested Extended, Main, and Diamond subsets scored by blockers and quality rubrics.

Metric: Main Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 16 models tracked.

Top models

#ModelScore
1Claude Sonnet 538.8
2Claude Opus 4.834.3
3Claude Opus 4.8 (xHigh)34.3
4GPT-5.525.5
5GPT-5.5 (xHigh)25.5
6Claude Opus 4.723
7GPT-5.4 Mini17.8
8Gemini 3.1 Pro (Preview)16.7
9Claude Sonnet 4.615.1
10MiniMax-M2.76
11MiniMax-M2.55.3
12Gemini 3.1 Flash Lite (Preview)4.8

Interactive version: theaggregate.ai/benchmark?slug=frontiercode · How the rankings work · Data refreshed daily, snapshot 2026-07-22.