HeaRTS - Feature Extraction: leaderboard

Metric: Mean task score (x100) over HeaRTS's feature extraction tasks (Perception category: deriving biomarkers or features from signals), each scored 0 to 1 by accuracy; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)82#54
2GLM-5 (Thinking)81#137 (GLM-5)
3Kimi K2 (Thinking)80#236 (Kimi K2)
4Qwen 3 Coder 480B A35B Instruct80#302
5GLM-4.7 (Thinking)80#185 (GLM-4.7)
6Claude Haiku 4.579#271
7DeepSeek V3.179#260
8Grok 4.1 Fast (Reasoning)78#208 (Grok 4.1 Fast)
9GPT-4.1 Mini77#346
10Gemini 2.5 Pro74#145
11Gemini 2.5 Flash74#237
12GPT-5 Mini74#176
13Llama 4 Maverick71#451
14MiniMax-M268#307
15Nemotron Nano 12B V249#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-feature-extraction · How It Works · Data refreshed daily, snapshot 2026-10-11.