MedConceal (Adaptive) - Confirmation Fine F1: leaderboard

Metric: Fine-grained F1 (0-1, times 100) of one-to-one matches between the submitted concerns and the gold concerns, credited only when supported by explicit patient-stated evidence in the dialogue, Task 1 (confirmation), an adaptive confirmation policy in which the model decides when to stop, capped at 20 turns, over MedConceal's 300 clinician-reviewed cases built from r/AskDocs threads, in dialogue with a reserved, stateful patient simulator that reveals hidden psychosocial concerns only under clinically meaningful elicitation; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.

Top models

#ModelScore
1GPT-5.251.2
2Claude Sonnet 4.550.9
3Qwen 3.5 9B7.9

Interactive version: theaggregate.ai/benchmark?slug=medconceal-adaptive-confirmation-fine-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.