SafePyramid - Exact Match: leaderboard

Metric: Rule matching rate (%) over all 3,000 SafePyramid cases: per case (a multi-turn conversation plus an in-context safety policy), the predicted set of violated rules matches the ground-truth set when neither false positives nor false negatives exceed (1 - tau) of the true violations; at tau = 1.0 only, i.e. the exact violated-rule set; per-policy evaluation, refused cases excluded; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 12 models tracked.

Top models

#ModelScore
1Claude Opus 4.7 (Max)35
2GPT-5.5 (xHigh)34.9
3DeepSeek V4 Pro (Max)31.7
4Seed 2.0 Pro (High)29.5
5Qwen 3.6 Max (Preview) (Thinking)29.2
6Gemini 3.5 Flash (High)28.1
7Hy3-preview (High)27.4
8Grok 4.3 (High)25.1
9gpt-oss-safeguard-20B13.5

Interactive version: theaggregate.ai/benchmark?slug=safepyramid-exact-match · How It Works · Data refreshed daily, snapshot 2026-09-29.