CRYSTAL - Answer Accuracy: leaderboard

Metric: Final-answer accuracy (%), fuzzy-matched against the gold answer, on the 6,372-question CRYSTAL test set, greedy decoding, JSON output of reasoning steps and answer (malformed outputs keep placeholder steps); higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 20 models tracked.

Top models

#ModelScoreOverall rank
1GPT-557.99#91
2GPT-5 Mini55.59#176
3Gemini 2.5 Flash53.95#237
4Qwen 2.5 VL 32B Instruct47.63#443
5GPT-5.2 Instant47.35#205
6Gemma 3 12B33.83#666
7Qwen 2.5 VL 7B Instruct30.43#643
8Gemma 3 4B28.65#1084
9MiniCPM-V-2.625.54#825
10Llama 3.2 11B24.83#1183

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=crystal-answer-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.