ConfBench - Expected Calibration Error: leaderboard

Metric: Expected calibration error of the verbalized confidence (0-1, lower is better; adaptive 5-bin quantile binning; verbalized top-4 guesses with probabilities in one call, OCR text with Amazon Textract confidences plus the document image, over the entity fields of 1,346 degraded variants of 75 RealKIE-FCC-Verified invoices). Source: arxiv.org. Saturation forecast: Estimated already saturated. 7 models tracked.

Top models

#ModelScore
1Claude Opus 4.60.05
2Claude Sonnet 4.50.13
3Claude Haiku 4.50.17
4Kimi K2.50.2
5Qwen 3.6 27B0.22
6Qwen 3 VL 235B A22B0.25
7Gemma 3 12B0.31

Interactive version: theaggregate.ai/benchmark?slug=confbench-expected-calibration-error · How It Works · Data refreshed daily, snapshot 2026-09-29.