HakushoBench (CoT): leaderboard

Metric: Accuracy (%), chain-of-thought prompt (think step by step, then answer); 2,053 manually written short-answer questions on chart and table images from 33 Japanese government white papers, judged correct or incorrect by GPT-5.1, mean of three runs, temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-5.1 (Medium)67.9
2GPT-4o (2024-11-20)54.1
3InternVL3.5-8B51

Interactive version: theaggregate.ai/benchmark?slug=hakushobench-cot · How It Works · Data refreshed daily, snapshot 2026-09-29.