AdversaRiskQA: leaderboard

Metric: Accuracy (%; mean of six domain-by-difficulty accuracies on the 225 prompts every model answered validly, GPT-5-mini judge). Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1Qwen 3 Next 80B A3B Instruct94.7
2GPT-591.4
3Qwen 3 30B A3B 2507 Instruct87.5
4Qwen 3 4B 2507 Instruct85.5
5GPT-OSS-120B76.8
6GPT-OSS-20B74.9

Interactive version: theaggregate.ai/benchmark?slug=adversariskqa · How It Works · Data refreshed daily, snapshot 2026-09-25.