ConfBench - Extraction Accuracy: leaderboard

Metric: Weighted overall accuracy (%; mean per-field similarity, normalized Levenshtein for strings and a tolerance comparator for numbers; verbalized top-4 guesses with probabilities in one call, OCR text with Amazon Textract confidences plus the document image, over the entity fields of 1,346 degraded variants of 75 RealKIE-FCC-Verified invoices). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.577
2Kimi K2.576
3Claude Opus 4.675
4Qwen 3.6 27B74
5Claude Haiku 4.571
6Gemma 3 12B65
7Qwen 3 VL 235B A22B64

Interactive version: theaggregate.ai/benchmark?slug=confbench-extraction-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.