SEAL - SWE Atlas - Codebase QnA — leaderboard

Scale SEAL evaluation of LLM ability to answer questions about large codebases, testing code comprehension and navigation.

Metric: Score. Source: scale.com. Status: saturation imminent. 18 models tracked.

Top models

#ModelScore
1GLM-5.248.12
2GPT-5.5 (xHigh)45.43
3GPT-5.4 (xHigh)36.3
4DeepSeek V4 Pro27.15
5Muse Spark24.2
6GLM-520.5
7Gemini 3.1 Pro (Preview)13.5
8Kimi K2.513.1
9MiniMax-M2.510.3
10Gemini 3 Flash8.2

Interactive version: theaggregate.ai/benchmark?slug=seal-swe-atlas-codebase-qna · How the rankings work · Data refreshed daily, snapshot 2026-07-22.