ClinTraceBench - Trend Classification: leaderboard

Metric: Accuracy (%; task T2, trend classification: direction and magnitude of a lab across two to five visits, 889 questions; full-context strategy: the whole multi-visit patient-clinician dialogue derived from MIMIC-IV records is given verbatim, models called through OpenRouter). Source: arxiv.org. Saturation forecast: Around September 2027. 4 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.647.4
2DeepSeek V3.137.5
3Claude Haiku 4.534.9
4GPT-4o Mini25.5

Interactive version: theaggregate.ai/benchmark?slug=clintracebench-trend-classification · How It Works · Data refreshed daily, snapshot 2026-09-26.