Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

HAL SWE-bench Verified Mini — leaderboard

Metric: Score (%). Source: www.swebench.com. 18 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (High)72
2Claude Sonnet 4.568
3Claude Opus 4.1 (20250805)61
4O4 Mini (Low)54
5Claude 3.7 Sonnet (High)54
6Claude Opus 4.1 (High)54
7Claude 3.7 Sonnet (20250219)50
8O4 Mini (High)50
9Claude Opus 4 (20250514)50
10GPT-5 (Medium)46
11O3 (Medium)46
12GPT-4.1 (2025-04-14)44
13Claude Haiku 4.5 (20251001)24
14DeepSeek V3 (0324)24
15Gemini 2.0 Flash (001)24

Interactive version: theaggregate.ai/benchmark?slug=hal-swe-bench-verified-mini · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.