Chronicles-OCR - Mature Text Parsing: leaderboard
Metric: Paragraph-level normalized edit similarity (0-1: one minus the edit distance over the longer sequence) of the transcription in reading order against the reference, on the mature scripts (Clerical, Regular, Running and Cursive), the paper Average column over the per-script results; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 29 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3.5 397B A17B (Non-reasoning) | 0.73 |
| 2 | Seed 2.0 Pro (Non-reasoning) | 0.72 |
| 3 | Seed 2.0 Pro | 0.71 |
| 4 | Kimi K2.5 (Non-reasoning) | 0.71 |
| 5 | Qwen 3.5 35B A3B (Non-reasoning) | 0.71 |
| 6 | Gemini 3.1 Pro (Preview) | 0.7 |
| 7 | Kimi K2.5 | 0.7 |
| 8 | Seed 1.8 | 0.67 |
| 9 | Qwen 3 VL 8B Instruct | 0.66 |
| 10 | Qwen 3 VL 235B A22B Instruct | 0.66 |
| 11 | Qwen 3 VL 235B A22B (Thinking) | 0.65 |
| 12 | MiMo-V2-Omni | 0.56 |
| 13 | Gemini 2.5 Pro | 0.53 |
| 14 | Claude Opus 4.7 (Thinking) | 0.5 |
| 15 | Qwen 2.5 VL 72B Instruct | 0.49 |
Interactive version: theaggregate.ai/benchmark?slug=chronicles-ocr-mature-text-parsing · How It Works · Data refreshed daily, snapshot 2026-10-07.