FinixDocBench - FinixInner: leaderboard
Metric: Overall score (0-100): mean over the five document categories of the per-category overall parsing score, the composite of text edit distance, table TEDS and reading-order edit distance; internal held-out track of 4,000 camera-captured documents from financial and insurance workflows: claim medical records, expense statements, expense receipts, medical examination reports and identity documents, 800 each. Source: arxiv.org. Saturation forecast: Around December 2026. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.5 | 76.91 |
| 2 | Qwen 3 VL 235B A22B Instruct | 76.32 |
| 3 | Qwen 3.5 397B A17B | 70.83 |
Interactive version: theaggregate.ai/benchmark?slug=finixdocbench-finixinner · How It Works · Data refreshed daily, snapshot 2026-09-29.