MedConceal (8-Turn) - Confirmation Coarse F1: leaderboard
Metric: Coarse-grained F1 (0-1, times 100) of the concern categories the model submits after the dialogue against the gold concerns' four categories, Task 1 (confirmation), a fixed 8-turn dialogue (matched to the mean human length), over MedConceal's 300 clinician-reviewed cases built from r/AskDocs threads, in dialogue with a reserved, stateful patient simulator that reveals hidden psychosocial concerns only under clinically meaningful elicitation; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 61.9 |
| 2 | GPT-5.2 | 55.4 |
| 3 | Qwen 3.5 9B | 22.5 |
Interactive version: theaggregate.ai/benchmark?slug=medconceal-8-turn-confirmation-coarse-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.