MatPhaseBench: leaderboard

Metric: BERTScore recall (baseline-rescaled XLNet-large BERTScore, at most 1 and unbounded below in principle, mean over the 200 diagrams; free-form description of each of the 200 literature phase diagrams, prompted zero-shot with its annotated semantic dimensions, compared with the expert description from the source paper; temperature 0). Source: arxiv.org. Saturation forecast: Around 2030. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.80.41
2GLM-5V Turbo0.4
3GPT-5.50.39
4Gemini 3.1 Pro (Preview)0.39
5Qwen 3.6 Plus0.39
6Qwen 3.6 27B0.38
7Qwen 3.6 35B A3B0.38
8Qwen 3.6 Flash0.38
9GLM-4.6V0.37
10Gemini 3.1 Flash Lite0.36
11Llama 4 Maverick0.3
12Llama 4 Scout0.3

Interactive version: theaggregate.ai/benchmark?slug=matphasebench · How It Works · Data refreshed daily, snapshot 2026-09-29.