SEAL - SWE Atlas - Test Writing — leaderboard

Scale SEAL evaluation of LLM ability to write comprehensive test suites for real-world software engineering tasks.

Metric: Score. Source: scale.com. Status: saturation imminent. 19 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)42.59
2GLM-5.241.48
3GPT-5.4 (xHigh)40
4Muse Spark31.11
5Gemini 3 Flash30.3
6Gemini 3.1 Pro (Preview)29.84
7GLM-528.74
8DeepSeek V4 Pro27.05
9Kimi K2.525.77
10MiniMax-M2.518.6

Interactive version: theaggregate.ai/benchmark?slug=seal-swe-atlas-test-writing · How the rankings work · Data refreshed daily, snapshot 2026-07-22.