ECG-Reasoning-Benchmark (PTB-XL) - Completion: leaderboard
Metric: Completion (%): share of samples whose every reasoning step (criterion selection, finding identification, grounding to the signal, diagnostic decision) is answered correctly, the chain stopping at the first error, on the 3,076 PTB-XL samples, multi-turn ECG reasoning chains for 17 core diagnoses (balanced positive and negative cases), 12-lead ECGs given as plotted images (OpenTSLM as a 100 Hz time series), temperature 0; each step's answer is checked for semantic agreement with the ground truth by a Gemini-3-Flash judge; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 19 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 6.26 | #78 |
| 2 | GPT-5.2 | 5.84 | #105 |
| 3 | Qwen 3 VL 32B Instruct | 5.71 | #276 |
| 4 | Qwen 3 VL 8B Instruct | 5.67 | #401 |
| 5 | Gemini 2.5 Flash | 5.12 | #237 |
| 6 | Llama 3.2 90B Vision Instruct | 4.47 | #611 |
| 7 | Gemini 2.5 Pro | 3.44 | #145 |
| 8 | GPT-5 Mini (2025-08-07) | 2.37 | #165 |
| 9 | MedGemma-4B-IT | 0.58 | #842 |
| 10 | Llama 3.2 11B Instruct | 0.49 | #1112 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=ecg-reasoning-benchmark-ptb-xl-completion · How It Works · Data refreshed daily, snapshot 2026-10-11.