ControBench - Trump: leaderboard

Metric: Macro-F1 (%; 2-class user stance, one-shot prompting on a stratified 200-user sample). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Kimi K289.46
2Qwen 3 235B A22B 2507 (Thinking)87.48
3DeepSeek R187.47
4DeepSeek V3 (0324)85.93
5Qwen 3 235B A22B84.45
6GPT-4o Mini74.97
7Llama 3.1 8B Instruct65.24

Interactive version: theaggregate.ai/benchmark?slug=controbench-trump · How It Works · Data refreshed daily, snapshot 2026-09-25.