AstroMind - Threat Assessment: leaderboard

Metric: Accuracy (%; 92 open-ended threat-level assessments checked by a DeepSeek-R1 judge against the reference; empty responses excluded). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.

Top models

#ModelScore
1QwQ-32B66.67
2GPT-OSS-20B59.09
3Qwen 3 32B49.38
4Gemma 3 27B46.74
5Qwen 3 8B45.45
6Gemma 3 4B26.09

Interactive version: theaggregate.ai/benchmark?slug=astromind-threat-assessment · How It Works · Data refreshed daily, snapshot 2026-09-26.