CRYSTAL: leaderboard

Metric: Match F1 (0-1, shown times 100): F1 of the model's reasoning steps against the reference steps, greedy one-to-one matching at cosine similarity 0.35 with all-distilroberta-v1, on the 6,372-question CRYSTAL test set, greedy decoding, JSON output of reasoning steps and answer (malformed outputs keep placeholder steps); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini77.3#176
2Gemini 2.5 Flash67.3#237
3Qwen 2.5 VL 32B Instruct65.3#443
4Gemma 3 4B61.8#1084
5GPT-561.2#91
6Gemma 3 12B60.5#666
7GPT-5.2 Instant56.4#205
8Qwen 2.5 VL 7B Instruct47.5#643
9Llama 3.2 11B47.1#1183
10MiniCPM-V-2.621.5#825

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=crystal · How It Works · Data refreshed daily, snapshot 2026-10-11.