PodBench - Thinking Mode - Script Quality: leaderboard

Metric: Podcast Script Quality (0-100 rubric: depth 45, structure 30, language 25). Source: arxiv.org. Saturation forecast: Estimated already saturated. 12 models tracked.

Top models

#ModelScore
1DeepSeek R1 052861.15
2Qwen 3 235B A22B (Thinking)60.42
3Qwen 3 32B (Thinking)59.92
4Qwen 3 14B (Reasoning)56.86
5Qwen 3 30B A3B (Thinking)54.78
6Qwen 3 8B (Thinking)53.89
7DeepSeek R1 0528 Qwen3 8B51.3
8Qwen 3 4B (Reasoning)51.15
9DeepSeek R1 Distill Qwen 32B48.05
10DeepSeek R1 Distill Llama 8B45.22
11Qwen 3 1.7B (Thinking)42.29
12DeepSeek-R1-Distill-Qwen-7B36.21

Interactive version: theaggregate.ai/benchmark?slug=podbench-thinking-mode-script-quality · How It Works · Data refreshed daily, snapshot 2026-09-25.