ClinTraceBench - Attribution Detection: leaderboard

Metric: Accuracy (%; task T3, controlled attribution detection: whether the dialogue attributes a finding to a problem, balanced yes/no, 600 questions; full-context strategy: the whole multi-visit patient-clinician dialogue derived from MIMIC-IV records is given verbatim, models called through OpenRouter). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1DeepSeek V3.193.8
2GPT-4o Mini93.1
3Claude Sonnet 4.688.5
4Claude Haiku 4.586.9

Interactive version: theaggregate.ai/benchmark?slug=clintracebench-attribution-detection · How It Works · Data refreshed daily, snapshot 2026-09-26.