PorTEXTO: leaderboard

Metric: Mean ANLS (%) across the four subsets (handwritten region and full page, synthetic, in-the-wild), a normalized edit-distance text-similarity score; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.

Top models

#ModelScore
1Qwen3.6-27B [28]86.8
2Gemma-4-31B-it [8]85.3
3Qwen3.5-9B [27]83.8
4Ministral3-Instruct-8B [13]77
5Gemma-4-E4B-it [8]74.8
6LLaVA-OneVision-1.5-8B [1]67.3
7Nemotron-Nano-12B-VL-V2 [5]65
8InternVL3.5-38B [32]63.9
9Bee-RL-8B [37]62.5
10AMALIA-VL-DPO [6]61.4

Interactive version: theaggregate.ai/benchmark?slug=portexto · How It Works · Data refreshed daily, snapshot 2026-09-29.