Chronicles-OCR - Mature Text Parsing: leaderboard

Metric: Paragraph-level normalized edit similarity (0-1: one minus the edit distance over the longer sequence) of the transcription in reading order against the reference, on the mature scripts (Clerical, Regular, Running and Cursive), the paper Average column over the per-script results; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 29 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B (Non-reasoning)0.73
2Seed 2.0 Pro (Non-reasoning)0.72
3Seed 2.0 Pro0.71
4Kimi K2.5 (Non-reasoning)0.71
5Qwen 3.5 35B A3B (Non-reasoning)0.71
6Gemini 3.1 Pro (Preview)0.7
7Kimi K2.50.7
8Seed 1.80.67
9Qwen 3 VL 8B Instruct0.66
10Qwen 3 VL 235B A22B Instruct0.66
11Qwen 3 VL 235B A22B (Thinking)0.65
12MiMo-V2-Omni0.56
13Gemini 2.5 Pro0.53
14Claude Opus 4.7 (Thinking)0.5
15Qwen 2.5 VL 72B Instruct0.49

Interactive version: theaggregate.ai/benchmark?slug=chronicles-ocr-mature-text-parsing · How It Works · Data refreshed daily, snapshot 2026-10-07.