HCSU: leaderboard

Metric: Average accuracy (%) over the stone-rubbing (Bei) and ink-manuscript (Tie) domains, 8-way controlled candidate selection: identify which of 8 calligraphers (1 correct, 7 distractors, each with a style description and reference image) wrote a target character image, temperature 0.1; chance 12.5; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1GPT-5.234.58
2GPT-4o32.93
3Claude Sonnet 4.532.82
4Qwen 3 VL 235B A22B Instruct30.25
5Gemini 2.5 Pro29.32
6Qwen 2.5 VL 72B Instruct24.93
7Qwen 3 VL 30B A3B Instruct21.81
8Qwen 2.5 VL 7B Instruct20.25

Interactive version: theaggregate.ai/benchmark?slug=hcsu · How It Works · Data refreshed daily, snapshot 2026-09-29.