SecIT Bench (Codex): leaderboard

The SecIT Bench scenario set run inside the Codex agent product rather than a plain model harness. Scored identically to the Pydantic AI board, and kept separate because the scaffold is part of what is being measured: GPT-5.6 Terra scores 71.9 here against 69.4 under the plain harness.

Metric: Accuracy (%). Source: secitbench.cribl.io. Status: years away from saturation. 3 models tracked.

Top models

#ModelScore
1GPT-5.6 Terra71.88
2GPT-5.6 Sol71.67
3GPT-5.6 Luna60.13

Interactive version: theaggregate.ai/benchmark?slug=secit-bench-codex · How It Works · Data refreshed daily, snapshot 2026-09-05.