RealVuln: leaderboard

Metric: Strict F3 score (recall weighted nine times over precision) of vulnerability findings against 796 hand-labeled entries (676 vulnerabilities, 120 false-positive traps) in 26 intentionally vulnerable Python repositories, agentic scanning with file, search and shell tools; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.651.7
2Gemini 3.1 Pro (Preview)51
3Claude Opus 4.647.7
4Kimi K2.546.6
5GLM-545.8
6MiniMax-M2.739
7Qwen 3.5 397B A17B38.1
8Claude Haiku 4.537.2
9Grok 4.20 (Reasoning)28.4
10Grok 321.3

Interactive version: theaggregate.ai/benchmark?slug=realvuln · How It Works · Data refreshed daily, snapshot 2026-10-07.