PorTEXTO: leaderboard
Metric: Mean ANLS (%) across the four subsets (handwritten region and full page, synthetic, in-the-wild), a normalized edit-distance text-similarity score; higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3.6-27B [28] | 86.8 |
| 2 | Gemma-4-31B-it [8] | 85.3 |
| 3 | Qwen3.5-9B [27] | 83.8 |
| 4 | Ministral3-Instruct-8B [13] | 77 |
| 5 | Gemma-4-E4B-it [8] | 74.8 |
| 6 | LLaVA-OneVision-1.5-8B [1] | 67.3 |
| 7 | Nemotron-Nano-12B-VL-V2 [5] | 65 |
| 8 | InternVL3.5-38B [32] | 63.9 |
| 9 | Bee-RL-8B [37] | 62.5 |
| 10 | AMALIA-VL-DPO [6] | 61.4 |
Interactive version: theaggregate.ai/benchmark?slug=portexto · How It Works · Data refreshed daily, snapshot 2026-09-29.