TOSSS (C/C++): leaderboard

Metric: TOSSS security score (0-1, times 100): the share of 500 C/C++ function pairs (the function before and after a CVE fix, mined from MegaVul) for which the model picks the secure version when shown both in random order with a neutral prompt that does not mention security; a constant or random choice scores about 50; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1GLM-587.8#137
2GPT-5.486.6#76
3Claude Opus 4.686.4#60
4Kimi K2.585.8#139
5MiniMax-M2.583.8#295
6Gemini 3 Flash (Preview)83.6#78
7Claude Sonnet 4.680.8#85
8Claude 3.5 Sonnet75.2#337
9Llama 3 70B Instruct71.9#624
10Gemini 3.1 Flash Lite (Preview)71#186
11codestral-250868#518
12DeepSeek V3.265.2#198
13Qwen3 Coder Next64.4#321
14Mistral Large 348#388

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tosss-c-c-plus-plus · How It Works · Data refreshed daily, snapshot 2026-10-11.