MTBBench-Longitudinal: leaderboard

Metric: Accuracy (%; unweighted mean of the three task accuracies; 40 patients from MSK-CHORD with clinical timelines and genomic files, 183 true/false and multiple-choice questions; multi-turn agent that requests patient files, no tools). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1GPT-4o (2024-08-06)64.2
2O4 Mini (2025-04-16)60

Interactive version: theaggregate.ai/benchmark?slug=mtbbench-longitudinal · How It Works · Data refreshed daily, snapshot 2026-09-26.