AstroMind - Threat Assessment: leaderboard
Metric: Accuracy (%; 92 open-ended threat-level assessments checked by a DeepSeek-R1 judge against the reference; empty responses excluded). Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | QwQ-32B | 66.67 |
| 2 | GPT-OSS-20B | 59.09 |
| 3 | Qwen 3 32B | 49.38 |
| 4 | Gemma 3 27B | 46.74 |
| 5 | Qwen 3 8B | 45.45 |
| 6 | Gemma 3 4B | 26.09 |
Interactive version: theaggregate.ai/benchmark?slug=astromind-threat-assessment · How It Works · Data refreshed daily, snapshot 2026-09-26.