HeaRTS - Cross-modal Translation: leaderboard

Metric: Mean task score (x100) over HeaRTS's cross-modal translation tasks (Generation category: synthesizing one signal modality from another), each scored 0 to 1 by 1 - sMAPE/2; the model reasons over the signal files by writing and running Python code in a CodeAct agent loop with a minimal package set; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)69#54
2Qwen 3 Coder 480B A35B Instruct66#302
3Gemini 2.5 Pro65#145
4GLM-5 (Thinking)65#137 (GLM-5)
5GPT-5 Mini64#176
6MiniMax-M264#307
7GLM-4.7 (Thinking)64#185 (GLM-4.7)
8Kimi K2 (Thinking)63#236 (Kimi K2)
9Grok 4.1 Fast (Reasoning)63#208 (Grok 4.1 Fast)
10Gemini 2.5 Flash61#237
11Claude Haiku 4.560#271
12DeepSeek V3.160#260
13GPT-4.1 Mini57#346
14Llama 4 Maverick54#451
15Nemotron Nano 12B V244#656

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=hearts-cross-modal-translation · How It Works · Data refreshed daily, snapshot 2026-10-11.