AA IFBench — leaderboard

Artificial Analysis independent evaluation of IFBench: instruction-following on 58 diverse out-of-domain constraints (294 questions).

Metric: Accuracy (%). Source: artificialanalysis.ai. Status: saturation imminent. 449 models tracked.

Top models

#ModelScore
1Grok 4.20 0309 (Reasoning)82.93
2MiniMax-M382.86
3Grok 4.3 (High)81.29
4Qwen 3.7 Max80.54
5Nemotron Cascade 2 30B A3B80.41
6MiMo-V2.5-Pro79.86
7Nova 2.0 Pro Preview (Low)79.59
8Nova 2.0 Pro Preview (Medium)79.05
9Qwen 3.7 Plus77.96
10GPT-5.2 Codex (xHigh)77.62
11Gemini 3.1 Flash Lite77.21
12Gemini 3.1 Pro (Preview)77.14
13Qwen 3.6 Max Preview76.6
14DeepSeek V4 Pro (Reasoning, Max Effort)76.46
15Gemini 3.5 Flash (High)76.33

Interactive version: theaggregate.ai/benchmark?slug=aa-ifbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.