OCRTurk - Raw Text: leaderboard
Metric: Raw-text normalized edit distance (0-1): Levenshtein distance between the model's output and the character-verified reference divided by the longer of the two lengths, averaged over the 180 Turkish pages after equations, tables and figure tags are extracted and heading marks removed; lower is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | OCRTurk PaddleOCR-VL (checkpoint unspecified) | 0.08 | |
| 2 | HunyuanOCR | 0.09 | |
| 3 | olmOCR-2 | 0.09 | |
| 4 | DeepSeek-OCR | 0.12 | |
| 5 | Docling | 0.13 | |
| 6 | Nanonets-OCR2-3B | 0.17 | |
| 7 | NVIDIA-Nemotron-Parse-v1.1 | 0.27 |
Interactive version: theaggregate.ai/benchmark?slug=ocrturk-raw-text · How It Works · Data refreshed daily, snapshot 2026-10-11.