HeaRTS - Signal Imputation: leaderboard

Metric: Mean task score (x100) over HeaRTS's signal imputation tasks (Generation category: reconstructing missing portions of a signal), each scored 0 to 1 by 1 - sMAPE/2; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)85#54
2GLM-5 (Thinking)84#137 (GLM-5)
3Grok 4.1 Fast (Reasoning)83#208 (Grok 4.1 Fast)
4Kimi K2 (Thinking)82#236 (Kimi K2)
5Qwen 3 Coder 480B A35B Instruct82#302
6GLM-4.7 (Thinking)82#185 (GLM-4.7)
7MiniMax-M281#307
8Claude Haiku 4.580#271
9Gemini 2.5 Pro79#145
10GPT-4.1 Mini79#346
11GPT-5 Mini79#176
12Gemini 2.5 Flash77#237
13DeepSeek V3.177#260
14Llama 4 Maverick75#451
15Nemotron Nano 12B V261#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-signal-imputation · How It Works · Data refreshed daily, snapshot 2026-10-11.