ConfBench - Expected Calibration Error: leaderboard
Metric: Expected calibration error of the verbalized confidence (0-1, lower is better; adaptive 5-bin quantile binning; verbalized top-4 guesses with probabilities in one call, OCR text with Amazon Textract confidences plus the document image, over the entity fields of 1,346 degraded variants of 75 RealKIE-FCC-Verified invoices). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 0.05 |
| 2 | Claude Sonnet 4.5 | 0.13 |
| 3 | Claude Haiku 4.5 | 0.17 |
| 4 | Kimi K2.5 | 0.2 |
| 5 | Qwen 3.6 27B | 0.22 |
| 6 | Qwen 3 VL 235B A22B | 0.25 |
| 7 | Gemma 3 12B | 0.31 |
Interactive version: theaggregate.ai/benchmark?slug=confbench-expected-calibration-error · How It Works · Data refreshed daily, snapshot 2026-09-29.