PAGE Bench - Final Result Score: leaderboard

Metric: Final Result Score (0-100; 0.3 point and command match + 0.3 task completion + 0.2 visual similarity + 0.2 geometric logic, the last three from a Gemini-2.5-Pro judge). Source: arxiv.org. Saturation forecast: Around May 2027. 17 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)16.31
2Qwen 3.6 Plus11.72
3GPT-5.49.96
4Claude Sonnet 4.68.65
5GLM-4.5V7.27
6Qwen 3 VL 8B3.42

Interactive version: theaggregate.ai/benchmark?slug=page-bench-final-result-score · How It Works · Data refreshed daily, snapshot 2026-09-25.