ECG-Reasoning-Benchmark (PTB-XL) - Completion: leaderboard

Metric: Completion (%): share of samples whose every reasoning step (criterion selection, finding identification, grounding to the signal, diagnostic decision) is answered correctly, the chain stopping at the first error, on the 3,076 PTB-XL samples, multi-turn ECG reasoning chains for 17 core diagnoses (balanced positive and negative cases), 12-lead ECGs given as plotted images (OpenTSLM as a 100 Hz time series), temperature 0; each step's answer is checked for semantic agreement with the ground truth by a Gemini-3-Flash judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 19 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash (Preview)6.26#78
2GPT-5.25.84#105
3Qwen 3 VL 32B Instruct5.71#276
4Qwen 3 VL 8B Instruct5.67#401
5Gemini 2.5 Flash5.12#237
6Llama 3.2 90B Vision Instruct4.47#611
7Gemini 2.5 Pro3.44#145
8GPT-5 Mini (2025-08-07)2.37#165
9MedGemma-4B-IT0.58#842
10Llama 3.2 11B Instruct0.49#1112

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=ecg-reasoning-benchmark-ptb-xl-completion · How It Works · Data refreshed daily, snapshot 2026-10-11.