SEAL - SWE Atlas - Test Writing: leaderboard

Scale SEAL evaluation of LLM ability to write test suites for software engineering tasks.

Metric: Score. Source: scale.com. Status: years away from saturation. 22 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (xHigh)45.9
2GPT-5.5 (xHigh)42.59
3GLM-5.241.48
4GPT-5.4 (xHigh)40
5Muse Spark31.11
6Gemini 3 Flash30.3
7Gemini 3.1 Pro (Preview)29.84
8GLM-528.74
9DeepSeek V4 Pro27.05
10Kimi K2.525.77
11MiniMax-M2.518.6

Interactive version: theaggregate.ai/benchmark?slug=seal-swe-atlas-test-writing · How It Works · Data refreshed daily, snapshot 2026-09-05.