KaliBench: leaderboard

Metric: Exact Correct (%), Unrestricted mode: share of the 5,000 verified queries whose generated Kali Linux command matches the reference after canonicalization (documented aliases and reordered optional flags allowed, positional order enforced), with only the query given; higher is better. Source: arxiv.org. Saturation forecast: Around June 2028. 24 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol61.68
2GPT-5.5 (xHigh)51.68
3Claude Opus 544.02
4GLM-5.241.3
5DeepSeek V3.2 (Thinking)33
6Qwen3 Coder Next26.2
7Qwen 2.5 72B Instruct25.4
8GPT-OSS-120B24
9Qwen 3 32B (Thinking)24
10Llama 3.3 70B Instruct22.6
11Qwen 3 8B (Thinking)20.9
12Mistral Small 3.220.7
13Qwen 3 32B (Non-reasoning)20.6
14GPT-OSS-20B19.6
15Gemma 3 27B (IT)18.9

Interactive version: theaggregate.ai/benchmark?slug=kalibench · How It Works · Data refreshed daily, snapshot 2026-10-04.