OCRTurk - Turkish Characters: leaderboard

Metric: Turkish character sensitivity (0-1): one minus the share of the Turkish-specific letters (c-cedilla, g-breve, dotless i, o-umlaut, s-cedilla, u-umlaut and their capitals) in the reference text that the output gets wrong, averaged over the 180 pages; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 7 models tracked.

Top models

#ModelScoreOverall rank
1HunyuanOCR0.88
2OCRTurk PaddleOCR-VL (checkpoint unspecified)0.82
3DeepSeek-OCR0.81
4olmOCR-20.8
5Nanonets-OCR2-3B0.79
6Docling0.71
7NVIDIA-Nemotron-Parse-v1.10.47

Interactive version: theaggregate.ai/benchmark?slug=ocrturk-turkish-characters · How It Works · Data refreshed daily, snapshot 2026-10-11.