TurkBench - Instruction Following: leaderboard

Metric: LLM-as-a-judge score (0-100). Source: huggingface.co. 38 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 Exp96.8
2DeepSeek V3.194.92
3Qwen 3 Next 80B A3B Instruct94.27
4GLM-5 FP894.08
5GPT-OSS-120B93.6
6Qwen 3.5 397B A17B93.42
7GLM-4.692.58
8Qwen 3 30B A3B 2507 Instruct92.51
9Qwen 3.5 35B A3B92.39
10Qwen 3.5 27B91.82
11Qwen 3.5 122B A10B91.54
12Step 3.5 Flash91.07
13MiMo-V2-Flash90.98
14MiniMax-M290.88
15MiniMax-M2.589.29

Interactive version: theaggregate.ai/benchmark?slug=turkbench-instruction-following · How It Works · Data refreshed daily, snapshot 2026-09-19.