BridgeBench Overall: leaderboard
BridgeBench V3 coding leaderboard: each model is rated 1-10 on nine axes (reasoning, front end, back end, one-shot, security, trust, laziness, speed, cost) and ranked by Overall, the mean of the nine. Laziness and cost are inverted so higher is better. Speed and cost are not capability measures, so cheap, fast models gain on Overall. It replaced the head-to-head arena ladder in September 2026.
Metric: Overall (mean of nine axes, 1-10). Source: www.bridgebench.ai. Saturation forecast: Around December 2027. 12 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-6 | 7.2 |
| 2 | Claude Fable 5.1 | 6.9 |
| 3 | Claude Fable 5 | 5.9 |
| 4 | GPT-5.6 Sol | 5.7 |
| 5 | Grok 4.6 | 5.6 |
| 6 | Claude Opus 5 | 5.1 |
| 7 | Kimi K3 | 5 |
| 8 | GLM-5.3 | 4.9 |
| 9 | GPT-5.6 Terra | 4.8 |
| 10 | DeepSeek V4 Pro | 4.7 |
| 11 | GLM-5.3 Flash | 4.6 |
| 12 | GPT-5.6 Luna | 4.5 |
Interactive version: theaggregate.ai/benchmark?slug=bridgebench-overall · How It Works · Data refreshed daily, snapshot 2026-09-23.