HeaRTS - Subject Profiling: leaderboard

Metric: Mean task score (x100) over HeaRTS's subject profiling tasks (Inference category: inferring subject-level attributes or conditions), each scored 0 to 1 by accuracy; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)57#54
2GPT-4.1 Mini56#346
3Llama 4 Maverick56#451
4Kimi K2 (Thinking)56#236 (Kimi K2)
5GLM-5 (Thinking)56#137 (GLM-5)
6GPT-5 Mini55#176
7DeepSeek V3.155#260
8Grok 4.1 Fast (Reasoning)55#208 (Grok 4.1 Fast)
9Qwen 3 Coder 480B A35B Instruct55#302
10GLM-4.7 (Thinking)55#185 (GLM-4.7)
11Gemini 2.5 Flash54#237
12Claude Haiku 4.552#271
13MiniMax-M252#307
14Gemini 2.5 Pro51#145
15Nemotron Nano 12B V250#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-subject-profiling · How It Works · Data refreshed daily, snapshot 2026-10-11.