DiagFlowBench - Correct Abstention: leaderboard

Metric: Correct abstention (%): share of injected off-procedure operator turns (coverage gaps, undocumented malfunctions, unrelated questions) on which the model declines instead of mapping to a real but inadequate step or fabricating one, classified by a claude-haiku-4-5 judge, on 1,676 multi-turn maintenance conversations over 50 industrial diagnostic flowcharts from a consumer manufacturer, procedure graph given as JSON in the system prompt, temperature 0, via OpenRouter; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Llama 3.3 70B81.3
2Nemotron 3 Super73.9
3Gemini 2.5 Flash72
4Llama 4 Maverick65.7
5Qwen 3 235B A22B (Thinking)64.9
6GPT-4o Mini58.3
7GPT-OSS-120B54.7
8Llama 4 Scout39.4
9Qwen 3 30B A3B (Thinking)27.4

Interactive version: theaggregate.ai/benchmark?slug=diagflowbench-correct-abstention · How It Works · Data refreshed daily, snapshot 2026-09-29.