KaliBench: leaderboard
Metric: Exact Correct (%), Unrestricted mode: share of the 5,000 verified queries whose generated Kali Linux command matches the reference after canonicalization (documented aliases and reordered optional flags allowed, positional order enforced), with only the query given; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 24 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol | 61.68 |
| 2 | GPT-5.5 (xHigh) | 51.68 |
| 3 | Claude Opus 5 | 44.02 |
| 4 | GLM-5.2 | 41.3 |
| 5 | DeepSeek V3.2 (Thinking) | 33 |
| 6 | Qwen3 Coder Next | 26.2 |
| 7 | Qwen 2.5 72B Instruct | 25.4 |
| 8 | GPT-OSS-120B | 24 |
| 9 | Qwen 3 32B (Thinking) | 24 |
| 10 | Llama 3.3 70B Instruct | 22.6 |
| 11 | Qwen 3 8B (Thinking) | 20.9 |
| 12 | Mistral Small 3.2 | 20.7 |
| 13 | Qwen 3 32B (Non-reasoning) | 20.6 |
| 14 | GPT-OSS-20B | 19.6 |
| 15 | Gemma 3 27B (IT) | 18.9 |
Interactive version: theaggregate.ai/benchmark?slug=kalibench · How It Works · Data refreshed daily, snapshot 2026-10-04.