TOSSS (C/C++, Security Hint): leaderboard
Metric: TOSSS security score (0-1, times 100): the share of 500 C/C++ function pairs (the function before and after a CVE fix, mined from MegaVul) for which the model picks the secure version when shown both in random order with a prompt that asks for the most secure implementation; a constant or random choice scores about 50; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Kimi K2.5 | 89 | #139 |
| 2 | Claude Opus 4.6 | 88.4 | #60 |
| 3 | GLM-5 | 88 | #137 |
| 4 | GPT-5.4 | 86.8 | #76 |
| 5 | Gemini 3 Flash (Preview) | 86.6 | #78 |
| 6 | MiniMax-M2.5 | 86 | #295 |
| 7 | Claude Sonnet 4.6 | 84 | #85 |
| 8 | Claude 3.5 Sonnet | 79.8 | #337 |
| 9 | Gemini 3.1 Flash Lite (Preview) | 77.8 | #186 |
| 10 | Llama 3 70B Instruct | 72.1 | #624 |
| 11 | Qwen3 Coder Next | 67.8 | #321 |
| 12 | DeepSeek V3.2 | 66 | #198 |
| 13 | codestral-2508 | 60.8 | #518 |
| 14 | Mistral Large 3 | 54.8 | #388 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=tosss-c-c-plus-plus-security-hint · How It Works · Data refreshed daily, snapshot 2026-10-11.