Probability Operator Inference: leaderboard
Metric: Overall accuracy (%; micro-average over all responses on 14,320 procedurally generated English yes or no prompts over fifteen inference templates with gradable epistemic modals (must, probably, might), varying question form, negation and surface content; zero-shot, temperature 0, first answer token parsed, uncertain responses counted as incorrect; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 29 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 | 75.6 |
| 2 | Claude Sonnet 4.6 | 75.1 |
| 3 | GPT-5.4 Mini | 65.6 |
Interactive version: theaggregate.ai/benchmark?slug=probability-operator-inference · How It Works · Data refreshed daily, snapshot 2026-09-29.