HeaRTS - Event Localization: leaderboard

Metric: Mean task score (x100) over HeaRTS's event localization tasks (Inference category: locating the time span of an event in a signal), each scored 0 to 1 by IoU; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Around February 2028. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)50#54
2GLM-5 (Thinking)40#137 (GLM-5)
3GLM-4.7 (Thinking)38#185 (GLM-4.7)
4Kimi K2 (Thinking)36#236 (Kimi K2)
5Grok 4.1 Fast (Reasoning)34#208 (Grok 4.1 Fast)
6Gemini 2.5 Pro32#145
7Claude Haiku 4.532#271
8Qwen 3 Coder 480B A35B Instruct32#302
9MiniMax-M231#307
10Llama 4 Maverick30#451
11GPT-4.1 Mini29#346
12DeepSeek V3.129#260
13Gemini 2.5 Flash27#237
14GPT-5 Mini23#176
15Nemotron Nano 12B V222#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-event-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.