DiagnosticIQ Pro: leaderboard

Metric: Macro accuracy (%) on DiagnosticIQ Pro, the 10-option variant of each question (mean of the 16 per-asset accuracies), zero-shot multiple choice, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 30 models tracked.

Top models

#ModelScore
1Claude Opus 4.659.81
2GPT-5.459.21
3Gemini 3.1 Pro (Preview)57.74
4Claude 3.7 Sonnet56.63
5Qwen 3.6 35B A3B55.61
6Gemma 4 26B A4B47.65
7Claude Sonnet 4.647.21
8Gemma 4 31B45.18
9Llama 4 Maverick42.65
10Granite 3.3 8B Instruct42.39
11DeepSeek V341.38
12GPT-540.69
13Llama 3.1 405B38.82
14Gemini 2.5 Pro37.51
15Llama 3.3 70B Instruct36.56

Interactive version: theaggregate.ai/benchmark?slug=diagnosticiq-pro · How It Works · Data refreshed daily, snapshot 2026-10-07.