HeaRTS - Physiological Classification: leaderboard

Metric: Mean task score (x100) over HeaRTS's physiological classification tasks (Inference category: classifying the physiological state of a signal segment), each scored 0 to 1 by accuracy; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 16 models tracked.

Top models

#ModelScoreOverall rank
1MiniMax-M248#307
2GLM-4.7 (Thinking)45#185 (GLM-4.7)
3Kimi K2 (Thinking)44#236 (Kimi K2)
4Grok 4.1 Fast (Reasoning)44#208 (Grok 4.1 Fast)
5GLM-5 (Thinking)44#137 (GLM-5)
6Gemini 3.1 Pro (Preview)43#54
7GPT-4.1 Mini43#346
8Claude Haiku 4.543#271
9DeepSeek V3.143#260
10Qwen 3 Coder 480B A35B Instruct43#302
11GPT-5 Mini42#176
12Gemini 2.5 Flash41#237
13Gemini 2.5 Pro40#145
14Llama 4 Maverick40#451
15Nemotron Nano 12B V234#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-physiological-classification · How It Works · Data refreshed daily, snapshot 2026-10-11.