AA IFBench: leaderboard
Artificial Analysis independent evaluation of IFBench: instruction-following on 58 out-of-domain constraints (294 questions).
Metric: Accuracy (%). Source: artificialanalysis.ai. Status: years away from saturation. 450 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.3 (Medium) | 83.33 |
| 2 | Grok 4.20 0309 (Reasoning) | 82.93 |
| 3 | MiniMax-M3 | 82.86 |
| 4 | Nemotron 3 Ultra 550B A55B (Reasoning) | 81.36 |
| 5 | Grok 4.3 (High) | 81.29 |
| 6 | Grok 4.20 0309 v2 (Reasoning) | 81.22 |
| 7 | Grok 4.3 (Low) | 80.95 |
| 8 | Qwen 3.7 Max | 80.54 |
| 9 | Nemotron Cascade 2 30B A3B | 80.41 |
| 10 | MiMo-V2.5-Pro | 79.86 |
| 11 | Nova 2.0 Pro Preview (Low) | 79.59 |
| 12 | DeepSeek V4 Flash (Reasoning, Max Effort) | 79.18 |
| 13 | Nova 2.0 Pro Preview (Medium) | 79.05 |
| 14 | Qwen 3.5 397B A17B (Reasoning) | 78.78 |
| 15 | Qwen 3.7 Plus | 77.96 |
Interactive version: theaggregate.ai/benchmark?slug=aa-ifbench · How It Works · Data refreshed daily, snapshot 2026-09-05.