Multi-IF — leaderboard

Multi-IF evaluates model capability on instruction following tasks from the linked upstream source with Average as the primary reported metric.

Metric: Average (self-reported). Source: benchmarklist.com. Status: saturation imminent. 14 models tracked.

Top models

#ModelScore
1O1 Preview87.7
2Llama 3.1 405B Instruct85.4
3O1 Mini85.3
4GPT-4o84.3
5Qwen 2.5 72B Instruct83.7
6Llama 3.1 70B82.6
7Claude 3.5 Sonnet81.7
8GPT-481.5
9Mistral Large 2 (Nov) Instruct (2411)80.5
10Claude 3 Sonnet78.2
11Gemini 1.5 Pro75.8
12Claude 3 Haiku72.9
13Gemini 1.5 Flash72.5
14Llama 3.1 8B68.8

Interactive version: theaggregate.ai/benchmark?slug=multi-if · How the rankings work · Data refreshed daily, snapshot 2026-07-22.