SafePyramid - L2 Novel Frameworks: leaderboard

Metric: Rule matching rate (%) on the 1,000 L2 cases (policies built on a novel regulatory framework defined in context): per case (a multi-turn conversation plus an in-context safety policy), the predicted set of violated rules matches the ground-truth set when neither false positives nor false negatives exceed (1 - tau) of the true violations; averaged over tau = 0.7, 0.8, 0.9 and 1.0; per-policy evaluation, refused cases excluded; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 12 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash (High)33.5
2GPT-5.5 (xHigh)32.9
3Seed 2.0 Pro (High)32.9
4Claude Opus 4.7 (Max)31
5Hy3-preview (High)30.9
6DeepSeek V4 Pro (Max)30.6
7Qwen 3.6 Max (Preview) (Thinking)26.4
8Grok 4.3 (High)11.9
9gpt-oss-safeguard-20B0.9

Interactive version: theaggregate.ai/benchmark?slug=safepyramid-l2-novel-frameworks · How It Works · Data refreshed daily, snapshot 2026-09-29.