CodeSpecBench-Func - Correctness: leaderboard

Metric: Correctness (%), the share of generated executable pre/postcondition specifications that accept every valid test, on CodeSpecBench-Func (2,494 LeetCode problems; 50 valid and 50 invalid inputs plus valid and model-generated buggy outputs per problem, validated by the online judge); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 21 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro88.1
2Claude Sonnet 4.5 (Thinking)85
3Gemini 2.5 Flash83.7
4QwQ-32B83.1
5GPT-OSS-120B82.6
6GPT-5 Mini82
7Qwen 3 32B81.9
8Qwen 3 14B80.1
9Qwen 3 8B75.5
10GPT-575.1
11GPT-OSS-20B73.7
12DeepSeek V3.2 (Non-reasoning)72.9
13Qwen 3 4B71.5
14Qwen 3 32B (Non-reasoning)65.8
15Qwen 3 14B (Non-reasoning)63.1

Interactive version: theaggregate.ai/benchmark?slug=codespecbench-func-correctness · How It Works · Data refreshed daily, snapshot 2026-10-07.