URSA Reaction Plausibility: leaderboard

Metric: Accuracy (%; binary plausible or implausible verdicts on the 1,000 expert-labeled reactions of URSA-reaction-plausibility-bench-2026, 500 plausible and 500 implausible, collected from synthesis-planner outputs; default reasoning effort; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 4 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)83
2GPT-5.576
3Grok 4.375
4Claude Opus 4.873

Interactive version: theaggregate.ai/benchmark?slug=ursa-reaction-plausibility · How It Works · Data refreshed daily, snapshot 2026-09-29.