CodeSpecBench-Repo - Correctness: leaderboard

Metric: Correctness (%), the share of generated executable pre/postcondition specifications that accept every valid test, on CodeSpecBench-Repo (500 SWE-bench Verified issues; specifications injected around the issue-relevant functions, trigger tests on the fixed and buggy versions with UTBoost-augmented tests); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 21 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.5 (Thinking)37.4
2Gemini 2.5 Pro30.8
3DeepSeek V3.2 (Non-reasoning)20.2
4GPT-5 Mini19.6
5GPT-518.8
6QwQ-32B10.6
7GPT-OSS-20B9.4
8Gemini 2.5 Flash9.2
9Qwen 3 32B8.8
10GPT-OSS-120B8.6
11Qwen 3 14B6
12Qwen 3 14B (Non-reasoning)4.8
13Qwen 3 32B (Non-reasoning)3.2
14Qwen 3 8B2
15Qwen 3 8B (Non-reasoning)1.6

Interactive version: theaggregate.ai/benchmark?slug=codespecbench-repo-correctness · How It Works · Data refreshed daily, snapshot 2026-10-07.