RxScribe Bench - Fabrications per Run: leaderboard

Metric: Fabrication events per run (lower is better; invented medicine rows or wrongly inferred active ingredients, clinical values and non-clinical fields asserted without support in the image, summed over 600 runs and divided by the run count; 200 handwritten Indian outpatient prescription images, three independent cold runs per image, image and JSON schema only, deterministic field-by-field scoring against human ground truth with no model judge). Source: arxiv.org. Saturation forecast: Around May 2028. 4 models tracked.

Top models

#ModelScore
1Muse Spark 1.22.28
2GPT-5.6 Sol2.95
3Gemini 3.1 Pro (Preview)3.06
4Claude Opus 53.69

Interactive version: theaggregate.ai/benchmark?slug=rxscribe-bench-fabrications-per-run · How It Works · Data refreshed daily, snapshot 2026-09-26.