AgentIF: leaderboard

Agent instruction-following benchmark measuring constraint and instruction success across vanilla, condition, example, formatting, semantic, and tool constraints.

Metric: Constraint Success Rate (self-reported). Source: benchmarklist.com. Status: saturated. 15 models tracked.

Top models

#ModelScore
1O1 Mini59.8
2GPT-4o58.5
3Qwen 3 32B58.4
4QwQ-32B58.1
5DeepSeek R157.9
6DeepSeek V356.7
7Claude 3.5 Sonnet56.6
8Llama 3.1 70B Instruct56.3
9DeepSeek R1 Distill Qwen 32B55.1
10DeepSeek R1 Distill Llama 70B55
11Llama 3.1 8B Instruct53.6
12Mistral 7B Instruct (v0.3)46.8

Interactive version: theaggregate.ai/benchmark?slug=agentif · How It Works · Data refreshed daily, snapshot 2026-09-05.