WildHandBench: leaderboard
Metric: Overall score (%; mean of 100 minus the text normalized edit distance, the table TEDS and the formula CDM over 500 real handwritten documents (free text, tables and formulas) in four languages and nine real-world scenarios such as notes, forms, whiteboards and exam sheets). Source: arxiv.org. Saturation forecast: Around April 2028. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 71.85 |
| 2 | Gemini 3.5 Flash | 68.25 |
| 3 | Kimi K3 | 67.81 |
| 4 | Qwen 3.5 Plus | 65.55 |
| 5 | Qwen 3 VL 235B A22B (Thinking) | 62.19 |
| 6 | Claude Opus 4.8 | 57.7 |
Interactive version: theaggregate.ai/benchmark?slug=wildhandbench · How It Works · Data refreshed daily, snapshot 2026-09-29.