SciConvBench - Inconsistency Resolution - PDEs - Grounded Resolution: leaderboard

Metric: Conversation-Grounded Resolution Rate (%): share of pdes inconsistency resolution cases (requests with planted conflicting information) whose issues are all resolved and grounded in the clarification dialogue rather than silently assumed, the assistant may ask one clarification question per turn of a simulated scientist (Claude Sonnet 4.6, answering only from the hidden reference specification) for at most 11 turns before producing the final task specification, guided scientist-mode system prompt, judged by Gemini 2.5 Pro against the annotated issues; higher is better. Source: arxiv.org. Saturation forecast: Around September 2028. 5 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro86.3
2Gemini 2.5 Flash71.2
3GPT-OSS-120B63
4GPT-5.242.5
5Claude Sonnet 4.60

Interactive version: theaggregate.ai/benchmark?slug=sciconvbench-inconsistency-resolution-pdes-grounded-resolution · How It Works · Data refreshed daily, snapshot 2026-10-07.