MTBBench-Longitudinal: leaderboard
Metric: Accuracy (%; unweighted mean of the three task accuracies; 40 patients from MSK-CHORD with clinical timelines and genomic files, 183 true/false and multiple-choice questions; multi-turn agent that requests patient files, no tools). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o (2024-08-06) | 64.2 |
| 2 | O4 Mini (2025-04-16) | 60 |
Interactive version: theaggregate.ai/benchmark?slug=mtbbench-longitudinal · How It Works · Data refreshed daily, snapshot 2026-09-26.