RxScribe Bench - Fabrications per Run: leaderboard
Metric: Fabrication events per run (lower is better; invented medicine rows or wrongly inferred active ingredients, clinical values and non-clinical fields asserted without support in the image, summed over 600 runs and divided by the run count; 200 handwritten Indian outpatient prescription images, three independent cold runs per image, image and JSON schema only, deterministic field-by-field scoring against human ground truth with no model judge). Source: arxiv.org. Saturation forecast: Around May 2028. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Muse Spark 1.2 | 2.28 |
| 2 | GPT-5.6 Sol | 2.95 |
| 3 | Gemini 3.1 Pro (Preview) | 3.06 |
| 4 | Claude Opus 5 | 3.69 |
Interactive version: theaggregate.ai/benchmark?slug=rxscribe-bench-fabrications-per-run · How It Works · Data refreshed daily, snapshot 2026-09-26.