TOSSS (C/C++, Security Hint): leaderboard

Metric: TOSSS security score (0-1, times 100): the share of 500 C/C++ function pairs (the function before and after a CVE fix, mined from MegaVul) for which the model picks the secure version when shown both in random order with a prompt that asks for the most secure implementation; a constant or random choice scores about 50; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScoreOverall rank
1Kimi K2.589#139
2Claude Opus 4.688.4#60
3GLM-588#137
4GPT-5.486.8#76
5Gemini 3 Flash (Preview)86.6#78
6MiniMax-M2.586#295
7Claude Sonnet 4.684#85
8Claude 3.5 Sonnet79.8#337
9Gemini 3.1 Flash Lite (Preview)77.8#186
10Llama 3 70B Instruct72.1#624
11Qwen3 Coder Next67.8#321
12DeepSeek V3.266#198
13codestral-250860.8#518
14Mistral Large 354.8#388

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=tosss-c-c-plus-plus-security-hint · How It Works · Data refreshed daily, snapshot 2026-10-11.