SWE-QA-Pro (Direct Answer): leaderboard

Metric: Overall score (5-50): sum of the five 1-10 judge scores, answering directly without repository access, on 260 questions from 26 long-tail repositories with executable environments, each answer scored 1-10 by a GPT-5 judge (three runs averaged) against a Claude Code reference answer checked by human annotators; questions were kept only when GPT-4o, Claude Sonnet 4.5 and Gemini 2.5 Pro could not answer them well this way; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.128.74#240
2Qwen 3 32B27.91#424
3Claude Sonnet 4.527.69#138
4DeepSeek V3.227.55#198
5Qwen 3 8B26.61#667
6GPT-4o26.58#333
7Gemini 2.5 Pro25.48#145
8Devstral Small 225.44#484
9Llama 3.3 70B Instruct24.32#520

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=swe-qa-pro-direct-answer · How It Works · Data refreshed daily, snapshot 2026-10-11.