DiagnosticIQ Verbose: leaderboard

Metric: Macro accuracy (%) on DiagnosticIQ Verbose, the core questions with each symbolic rule paraphrased into natural language (mean of the per-asset accuracies), zero-shot multiple choice, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1Claude Opus 4.671.58
2Claude 3.7 Sonnet68.22
3Claude Sonnet 4.664.81
4O164.53
5DeepSeek V362.02
6Qwen 2.5 72B Instruct56.73
7Mistral Medium 355.89
8Gemini 1.5 Pro55.78
9Llama 3.3 70B Instruct55.16
10Llama 3.1 405B53.81
11Mistral Small 3.153.78
12Granite 3.3 8B Instruct53.69
13Gemini 2.0 Flash49.78
14Phi-445.64
15Claude 3.5 Haiku42.3

Interactive version: theaggregate.ai/benchmark?slug=diagnosticiq-verbose · How It Works · Data refreshed daily, snapshot 2026-10-07.