AA IFBench: leaderboard

Artificial Analysis independent evaluation of IFBench: instruction-following on 58 out-of-domain constraints (294 questions).

Metric: Accuracy (%). Source: artificialanalysis.ai. Status: years away from saturation. 450 models tracked.

Top models

#ModelScore
1Grok 4.3 (Medium)83.33
2Grok 4.20 0309 (Reasoning)82.93
3MiniMax-M382.86
4Nemotron 3 Ultra 550B A55B (Reasoning)81.36
5Grok 4.3 (High)81.29
6Grok 4.20 0309 v2 (Reasoning)81.22
7Grok 4.3 (Low)80.95
8Qwen 3.7 Max80.54
9Nemotron Cascade 2 30B A3B80.41
10MiMo-V2.5-Pro79.86
11Nova 2.0 Pro Preview (Low)79.59
12DeepSeek V4 Flash (Reasoning, Max Effort)79.18
13Nova 2.0 Pro Preview (Medium)79.05
14Qwen 3.5 397B A17B (Reasoning)78.78
15Qwen 3.7 Plus77.96

Interactive version: theaggregate.ai/benchmark?slug=aa-ifbench · How It Works · Data refreshed daily, snapshot 2026-09-05.