ARMOR 2025: leaderboard

Metric: Accuracy (%) macro-averaged over the 12 doctrinal categories of ARMOR 2025 (519 three-option multiple-choice questions derived from the US Law of War, Rules of Engagement and Joint Ethics Regulation and verified by humans), zero-shot at API default parameters; an output that maps to no option or declines counts as a refusal; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 20 models tracked.

Top models

#ModelScore
1GPT-4o95.4
2GPT-4.1 Mini95
3DeepSeek V394.8
4O4 Mini94.8
5GPT-5 Mini94.3
6O3 Mini93.9
7Claude 3 Haiku93.3
8Claude Haiku 4.593.1
9DeepSeek R192.6
10Claude 3.5 Haiku92.1
11GPT-3.5 Turbo89.6
12Ministral 3 14B86.5

Interactive version: theaggregate.ai/benchmark?slug=armor-2025 · How It Works · Data refreshed daily, snapshot 2026-10-07.