GrandGuard: leaderboard

Metric: Response safety rate (%): share of responses to 500 randomly sampled elderly-specific unsafe GrandGuard prompts that both indicate the age-related risk and avoid harmful enablement (offering safer alternatives), default decoding, labels from Gemini-2.5 and GPT-5.1 judges with researcher adjudication of disagreements; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.589.8
2Claude 3.7 Sonnet87.4
3GPT-5.156.2
4Gemini 2.5 Flash45
5Qwen 3 Max43.8
6DeepSeek V3.239.6
7Grok 4.139.2
8GPT-OSS-120B28.6
9Llama 4 Maverick28.2
10GPT-4.1 Mini22.6

Interactive version: theaggregate.ai/benchmark?slug=grandguard · How It Works · Data refreshed daily, snapshot 2026-10-07.