FullStackBench en — leaderboard
English subset of FullStackBench for evaluating end-to-end software engineering and full-stack development capability.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 28 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O1 Preview | 65.62 |
| 2 | O1 Mini | 64.73 |
| 3 | Claude 3.5 Sonnet | 62.57 |
| 4 | GPT-4o | 61.77 |
| 5 | DeepSeek V2.5 | 58.65 |
| 6 | Qwen 2.5 72B Instruct | 56.88 |
| 7 | Qwen 2.5 Coder 32B Instruct | 56.88 |
| 8 | Qwen 2.5 Coder 14B Instruct | 55.28 |
| 9 | Llama 3.1 70B Instruct | 51.45 |
| 10 | Qwen 2.5 Coder 7B Instruct | 47.95 |
| 11 | Yi-Coder-9B-Chat | 47.13 |
| 12 | OpenCoder-8B-Instruct | 43.63 |
| 13 | deepseek-coder-6.7B-instruct | 41.88 |
| 14 | OpenCoder-1.5B-Instruct | 33.52 |
| 15 | CodeLlama-34B-Instruct | 29.22 |
Interactive version: theaggregate.ai/benchmark?slug=fullstackbench-en · How the rankings work · Data refreshed daily, snapshot 2026-07-22.