SecIT Bench (Claude Code): leaderboard

The SecIT Bench scenario set run inside the Claude Code agent product rather than a plain model harness. Scored identically to the Pydantic AI board, and kept separate because the scaffold is part of what is being measured: the same model moves by up to three points between harnesses.

Metric: Accuracy (%). Source: secitbench.cribl.io. Status: years away from saturation. 3 models tracked.

Top models

#ModelScore
1Claude Opus 578.64
2Claude Opus 4.867.14
3Claude Sonnet 565.93

Interactive version: theaggregate.ai/benchmark?slug=secit-bench-claude-code · How It Works · Data refreshed daily, snapshot 2026-09-05.