AA IFBench — leaderboard
Artificial Analysis independent evaluation of IFBench: instruction-following on 58 diverse out-of-domain constraints (294 questions).
Metric: Accuracy (%). Source: artificialanalysis.ai. Status: saturation imminent. 449 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4.20 0309 (Reasoning) | 82.93 |
| 2 | MiniMax-M3 | 82.86 |
| 3 | Grok 4.3 (High) | 81.29 |
| 4 | Qwen 3.7 Max | 80.54 |
| 5 | Nemotron Cascade 2 30B A3B | 80.41 |
| 6 | MiMo-V2.5-Pro | 79.86 |
| 7 | Nova 2.0 Pro Preview (Low) | 79.59 |
| 8 | Nova 2.0 Pro Preview (Medium) | 79.05 |
| 9 | Qwen 3.7 Plus | 77.96 |
| 10 | GPT-5.2 Codex (xHigh) | 77.62 |
| 11 | Gemini 3.1 Flash Lite | 77.21 |
| 12 | Gemini 3.1 Pro (Preview) | 77.14 |
| 13 | Qwen 3.6 Max Preview | 76.6 |
| 14 | DeepSeek V4 Pro (Reasoning, Max Effort) | 76.46 |
| 15 | Gemini 3.5 Flash (High) | 76.33 |
Interactive version: theaggregate.ai/benchmark?slug=aa-ifbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.