Chronicles-OCR - Archaic Text Parsing: leaderboard
Metric: Paragraph-level normalized edit similarity (0-1: one minus the edit distance over the longer sequence) of the transcription in reading order against the reference, on the archaic scripts (Oracle Bone, Bronze and Seal), the paper Average column over the per-script results; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 29 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Kimi K2.5 | 0.22 |
| 2 | Kimi K2.5 (Non-reasoning) | 0.22 |
| 3 | Qwen 3.5 397B A17B (Non-reasoning) | 0.22 |
| 4 | Seed 2.0 Pro | 0.21 |
| 5 | Qwen 3.5 35B A3B (Non-reasoning) | 0.2 |
| 6 | Qwen 3 VL 235B A22B Instruct | 0.19 |
| 7 | Qwen 3 VL 8B Instruct | 0.18 |
| 8 | Seed 2.0 Pro (Non-reasoning) | 0.18 |
| 9 | Qwen 3 VL 235B A22B (Thinking) | 0.17 |
| 10 | Seed 1.8 | 0.17 |
| 11 | Gemini 3.1 Pro (Preview) | 0.15 |
| 12 | Qwen 3 VL 8B (Thinking) | 0.09 |
| 13 | MiMo-V2-Omni | 0.08 |
| 14 | Claude Opus 4.7 (Thinking) | 0.08 |
| 15 | Gemini 2.5 Pro | 0.07 |
Interactive version: theaggregate.ai/benchmark?slug=chronicles-ocr-archaic-text-parsing · How It Works · Data refreshed daily, snapshot 2026-10-07.