SWE Atlas - Codebase QnA: leaderboard

SWE Atlas Codebase QnA evaluates LLMs on deep code comprehension and question answering across real-world software repositories.

Metric: Score. Source: scale.com. Status: years away from saturation. 22 models tracked.

Top models

#ModelScore
1GLM-5.248.12
2GPT-5.6 Sol (xHigh)46
3GPT-5.5 (xHigh)45.43
4GPT-5.4 (xHigh)36.3
5DeepSeek V4 Pro27.15
6Muse Spark24.2
7GLM-520.5
8Gemini 3.1 Pro (Preview)13.5
9Kimi K2.513.1
10MiniMax-M2.510.3
11Gemini 3 Flash8.2

Interactive version: theaggregate.ai/benchmark?slug=swe-atlas-codebase-qna · How It Works · Data refreshed daily, snapshot 2026-09-05.