HeaRTS - Trajectory Analysis: leaderboard

Metric: Mean task score (x100) over HeaRTS's trajectory analysis tasks (Deduction category: analyzing health trajectories across visits or sessions), each scored 0 to 1 by accuracy; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 16 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.1 Mini55#346
2MiniMax-M255#307
3Nemotron Nano 12B V254#656
4DeepSeek V3.152#260
5GLM-5 (Thinking)52#137 (GLM-5)
6Claude Haiku 4.550#271
7Llama 4 Maverick49#451
8GPT-5 Mini48#176
9Qwen 3 Coder 480B A35B Instruct47#302
10GLM-4.7 (Thinking)47#185 (GLM-4.7)
11Gemini 2.5 Flash46#237
12Gemini 2.5 Pro45#145
13Grok 4.1 Fast (Reasoning)45#208 (Grok 4.1 Fast)
14Kimi K2 (Thinking)43#236 (Kimi K2)
15Gemini 3.1 Pro (Preview)39#54

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-trajectory-analysis · How It Works · Data refreshed daily, snapshot 2026-10-11.