PodBench - Thinking Mode - Instruction Following: leaderboard

Metric: Instruction Following (0-100; judge-derived checklist, items scored 0/0.5/1). Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScore
1Qwen 3 235B A22B (Thinking)94.32
2DeepSeek R1 052894.27
3Qwen 3 32B (Thinking)93.78
4Qwen 3 14B (Reasoning)91.79
5Qwen 3 30B A3B (Thinking)89.88
6Qwen 3 8B (Thinking)88.45
7DeepSeek R1 0528 Qwen3 8B85.67
8Qwen 3 4B (Reasoning)83.04
9DeepSeek R1 Distill Qwen 32B80.12
10DeepSeek R1 Distill Llama 8B66.94
11Qwen 3 1.7B (Thinking)62.65
12DeepSeek-R1-Distill-Qwen-7B42.59

Interactive version: theaggregate.ai/benchmark?slug=podbench-thinking-mode-instruction-following · How It Works · Data refreshed daily, snapshot 2026-09-25.