IFMTBench - Multi-Constraint: leaderboard

Metric: IF_Score (%): hard constraints checked by deterministic verifiers gate the item, soft constraints are scored by a gpt-oss-120b rubric judge, mean over items, on the 2,838 multi-constraint items, translation requests into seven languages with instructions paraphrased in all seven, officially recommended decoding with non-thinking mode for open models; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)84.53
2Qwen 3.5 27B (Non-reasoning)78.81
3Hy-MT2-30B-A3B75.8
4Gemma 4 26B A4B73.89
5Gemma 4 31B72.64
6Qwen 3.6 35B A3B (Non-reasoning)71.43
7Qwen 3.5 35B A3B (Non-reasoning)70.32
8Gemma 4 E4B69.25
9Qwen 3.5 9B (Non-reasoning)64.28
10Hy-MT2-1.8B57.61
11Qwen 3.5 4B (Non-reasoning)57.23
12Gemma 4 E2B50.72
13Qwen 3.5 2B (Non-reasoning)24.6
14Qwen 3.5 0.8B (Non-reasoning)7.46

Interactive version: theaggregate.ai/benchmark?slug=ifmtbench-multi-constraint · How It Works · Data refreshed daily, snapshot 2026-10-07.