ReCon: leaderboard
Metric: Balanced accuracy (%) of CONFORMITY or NON-CONFORMITY verdicts against certified-auditor assessments for the 93 ISO/IEC 27002:2022 controls on three real policy sets, over the study retrieval configurations (overall figure of Table 5; 50% is chance); retrieval-augmented prompts, local CPU inference at temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | ReCon granite4.1 8B (Ollama build) | 75 |
| 2 | ReCon granite4.1 3B (Ollama build) | 73.4 |
| 3 | ReCon granite3.3 8B (Ollama build) | 68 |
| 4 | ReCon llama3.2 3B (Ollama build) | 64.4 |
| 5 | ReCon phi4 14B (Ollama build) | 59.9 |
| 6 | ReCon mistral 7B (Ollama build) | 54.6 |
| 7 | ReCon gemma3 4B (Ollama build) | 54.3 |
| 8 | ReCon gemma3n E2B (Ollama build) | 52.4 |
Interactive version: theaggregate.ai/benchmark?slug=recon · How It Works · Data refreshed daily, snapshot 2026-09-29.