Chronicles-OCR - Archaic Text Parsing: leaderboard

Metric: Paragraph-level normalized edit similarity (0-1: one minus the edit distance over the longer sequence) of the transcription in reading order against the reference, on the archaic scripts (Oracle Bone, Bronze and Seal), the paper Average column over the per-script results; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 29 models tracked.

Top models

#ModelScore
1Kimi K2.50.22
2Kimi K2.5 (Non-reasoning)0.22
3Qwen 3.5 397B A17B (Non-reasoning)0.22
4Seed 2.0 Pro0.21
5Qwen 3.5 35B A3B (Non-reasoning)0.2
6Qwen 3 VL 235B A22B Instruct0.19
7Qwen 3 VL 8B Instruct0.18
8Seed 2.0 Pro (Non-reasoning)0.18
9Qwen 3 VL 235B A22B (Thinking)0.17
10Seed 1.80.17
11Gemini 3.1 Pro (Preview)0.15
12Qwen 3 VL 8B (Thinking)0.09
13MiMo-V2-Omni0.08
14Claude Opus 4.7 (Thinking)0.08
15Gemini 2.5 Pro0.07

Interactive version: theaggregate.ai/benchmark?slug=chronicles-ocr-archaic-text-parsing · How It Works · Data refreshed daily, snapshot 2026-10-07.