HG-Bench - Unified Score: leaderboard
Metric: Unified score (0-100): per-sample weighted combination of question-level and step-level page F1 averaged over all 500 samples, unparseable outputs scoring 0; 500 human-annotated multi-page handwritten K-12 homework samples; the model outputs page-aware question-level answer boxes and ordered step boxes as JSON with normalized 0-1000 coordinates, one shared prompt, default decoding, one format-reminder retry; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro (Preview) | 42.33 |
| 2 | GLM-5V Turbo | 40.1 |
| 3 | Qwen 3.5 397B A17B | 32.73 |
| 4 | Doubao-Seed-2.0-Pro-260215 | 21.22 |
| 5 | Kimi K2.5 | 20.18 |
| 6 | Claude Sonnet 4.6 | 8.76 |
| 7 | GPT-5.4 | 8.12 |
Interactive version: theaggregate.ai/benchmark?slug=hg-bench-unified-score · How It Works · Data refreshed daily, snapshot 2026-09-29.