AgentIF — leaderboard

Agent instruction-following benchmark measuring constraint and instruction success across vanilla, condition, example, formatting, semantic, and tool constraints.

Metric: Constraint Success Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1O1 Mini59.8
2GPT-4o58.5
3Qwen 3 32B58.4
4QwQ-32B58.1
5DeepSeek V356.7
6Claude 3.5 Sonnet56.6
7Llama 3.1 70B Instruct56.3
8DeepSeek R1 Distill Llama 70B55
9Llama 3.1 8B Instruct53.6
10Mistral 7B Instruct (v0.3)46.8

Interactive version: theaggregate.ai/benchmark?slug=agentif · How the rankings work · Data refreshed daily, snapshot 2026-07-22.