PICQ-drama (Multi-Task Prompting): leaderboard
Metric: Fidelity F1 (%; up to five identified personas per question matched to the annotators' missing personas by category and NLI, multi-task prompt). Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 32B | 48 |
| 2 | Llama 3.1 8B Instruct | 41.8 |
| 3 | Qwen 3 8B | 39.1 |
| 4 | Llama 3.1 70B Instruct | 34.6 |
| 5 | GPT-4.1 | 30 |
Interactive version: theaggregate.ai/benchmark?slug=picq-drama-multi-task-prompting · How It Works · Data refreshed daily, snapshot 2026-09-25.