PICQ-drama: leaderboard

Metric: Fidelity F1 (%; up to five identified personas per question matched to the annotators' missing personas by category and NLI, baseline prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScore
1Qwen 3 32B46.8
2Llama 3.1 8B Instruct39.9
3Qwen 3 8B39.3
4GPT-4.137.4
5Llama 3.1 70B Instruct30.9

Interactive version: theaggregate.ai/benchmark?slug=picq-drama · How It Works · Data refreshed daily, snapshot 2026-09-25.