LongMedBench - Joint Sorting: leaderboard

Metric: Mean Kendall tau (-1 to 1; agreement between the ordering given by the model and the true chronological order, averaged over 2,878 items; joint sorting: pair ten admission and discharge fragments of five visits and order them in time; MIMIC-IV patients with long multi-visit event streams; naive long-context agent, provider default configuration). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1DeepSeek V3.2 (Thinking)0.33
2GPT-5 Mini0.3
3DeepSeek V3.2 (Non-reasoning)0.13
4Qwen Turbo0.03

Interactive version: theaggregate.ai/benchmark?slug=longmedbench-joint-sorting · How It Works · Data refreshed daily, snapshot 2026-09-29.