ThermoQA: leaderboard

Metric: Composite score (%): question-weighted average of the three tiers (110 property lookups, 101 component analyses, 82 cycle analyses) over the 293 open-ended engineering thermodynamics problems of ThermoQA (ground truth computed with CoolProp 7.2.0 for water, R-134a and variable-cp air), zero-shot with a symbol = value unit answer format, each model at its stated reasoning setting, mean of three independent runs; values are extracted by regex and gpt-4.1-mini and a property counts as correct within 2 percent relative or 0.5 absolute tolerance (0.03 for quality, 0.02 for efficiencies and COP); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 6 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)94.1
2GPT-5.4 (High)93.1
3Gemini 3.1 Pro (Preview)92.5
4MiniMax-M2.573

Interactive version: theaggregate.ai/benchmark?slug=thermoqa · How It Works · Data refreshed daily, snapshot 2026-10-07.