SciConvBench - Inconsistency Resolution - Fluid Mechanics - Grounded Resolution: leaderboard
Metric: Conversation-Grounded Resolution Rate (%): share of fluid mechanics inconsistency resolution cases (requests with planted conflicting information) whose issues are all resolved and grounded in the clarification dialogue rather than silently assumed, the assistant may ask one clarification question per turn of a simulated scientist (Claude Sonnet 4.6, answering only from the hidden reference specification) for at most 11 turns before producing the final task specification, guided scientist-mode system prompt, judged by Gemini 2.5 Pro against the annotated issues; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 65.2 |
| 2 | GPT-5.2 | 51 |
| 3 | Claude Sonnet 4.6 | 46.5 |
| 4 | Gemini 2.5 Flash | 45.8 |
| 5 | GPT-OSS-120B | 17.4 |
Interactive version: theaggregate.ai/benchmark?slug=sciconvbench-inconsistency-resolution-fluid-mechanics-grounded-resolution · How It Works · Data refreshed daily, snapshot 2026-10-07.