MedConceal (Adaptive) - Confirmation Fine F1: leaderboard
Metric: Fine-grained F1 (0-1, times 100) of one-to-one matches between the submitted concerns and the gold concerns, credited only when supported by explicit patient-stated evidence in the dialogue, Task 1 (confirmation), an adaptive confirmation policy in which the model decides when to stop, capped at 20 turns, over MedConceal's 300 clinician-reviewed cases built from r/AskDocs threads, in dialogue with a reserved, stateful patient simulator that reveals hidden psychosocial concerns only under clinically meaningful elicitation; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 51.2 |
| 2 | Claude Sonnet 4.5 | 50.9 |
| 3 | Qwen 3.5 9B | 7.9 |
Interactive version: theaggregate.ai/benchmark?slug=medconceal-adaptive-confirmation-fine-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.