Probability Operator Inference: leaderboard

Metric: Overall accuracy (%; micro-average over all responses on 14,320 procedurally generated English yes or no prompts over fifteen inference templates with gradable epistemic modals (must, probably, might), varying question form, negation and surface content; zero-shot, temperature 0, first answer token parsed, uncertain responses counted as incorrect; higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 29 models tracked.

Top models

#ModelScore
1Claude Opus 4.775.6
2Claude Sonnet 4.675.1
3GPT-5.4 Mini65.6

Interactive version: theaggregate.ai/benchmark?slug=probability-operator-inference · How It Works · Data refreshed daily, snapshot 2026-09-29.