SecCodeBench — leaderboard

Security benchmark for AI-generated and AI-repaired code, reporting secure-code repair and generation scores with and without hints.

Metric: Total Score. Source: alibaba.github.io. Status: saturation imminent. 36 models tracked.

Top models

#ModelScore
1Claude Opus 4.568.1
2Qwen 3.5 Plus (Thinking)67.08
3Claude Opus 4.664.9
4Qwen3 Coder Next62.76
5Gemini 3 Pro62.42
6GLM-5 (Thinking)62.13
7Kimi K2.5 (Thinking)61.25
8GPT-5.459.74
9Gemini 3 Flash58.66
10GPT-5.258.23
11Claude Sonnet 4.556.83
12DeepSeek V3.2 (Thinking)55.24
13Kimi K2.5 (Non-reasoning)55.22
14Gemini 3.1 Pro (Preview)55.21
15DeepSeek R1 052854.06

Interactive version: theaggregate.ai/benchmark?slug=seccodebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.