HeaRTS: leaderboard

Metric: Overall score (x100): macro-average of the Perception, Inference, Generation and Deduction category scores over HeaRTS's 110 health time-series tasks (16 datasets, 20,226 test cases); each task is scored 0 to 1 by accuracy, IoU or 1 - sMAPE/2 on min-max normalized signals; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)69#54
2GLM-5 (Thinking)68#137 (GLM-5)
3GLM-4.7 (Thinking)66#185 (GLM-4.7)
4Kimi K2 (Thinking)65#236 (Kimi K2)
5Grok 4.1 Fast (Reasoning)65#208 (Grok 4.1 Fast)
6DeepSeek V3.164#260
7GPT-4.1 Mini63#346
8Claude Haiku 4.563#271
9Qwen 3 Coder 480B A35B Instruct63#302
10Gemini 2.5 Pro62#145
11Gemini 2.5 Flash62#237
12GPT-5 Mini62#176
13MiniMax-M261#307
14Llama 4 Maverick58#451
15Nemotron Nano 12B V247#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts · How It Works · Data refreshed daily, snapshot 2026-10-11.