MedErrBench (English) - Error Localization: leaderboard

Metric: Localization accuracy (%): naming the sentence that contains the error (an error-free note must be answered 'NA', which counts as correct when the gold label is NA); 208 English test notes adapted from MedQA cases (104 with an error, 104 without) with clinician-verified inserted errors; main-table prompting (the paper does not state its shot setting); printed as fractions and shown times 100; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 14 models tracked.

Top models

#ModelScoreOverall rank
1Doubao-1.5-Thinking-Pro77.4#133
2GPT-4o Mini52.4#588
3Qwen 2.5 7B Instruct49#846
4MedGemma-4B43.8#731
5Llama 3.1 8B Instruct36.1#1018
6GPT-4o34.6#333
7Gemini 2.5 Flash Lite26.4#413
8Llama 3.3 70B Instruct25.5#520
9Gemini 2.0 Flash16.8#331

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mederrbench-english-error-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.