Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

Senior SWE-Bench — leaderboard

Metric: Tasteful Solve Rate (pass@1, %). Source: senior-swe-bench.snorkel.ai. 10 models tracked.

Top models

#ModelScore
1Claude Fable 529.1
2Claude Opus 4.825
3GPT-5.6 Sol24.4
4Claude Sonnet 517.4
5Grok 4.517.2
6GPT-5.515.9
7Claude Opus 4.713.8
8GPT-5.413.6
9GLM-5.213.1
10Kimi K2.69.4

Interactive version: theaggregate.ai/benchmark?slug=senior-swe-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.