SABRE-Prior: leaderboard
Metric: Macro-average accuracy (%; unweighted mean of the four primary subset scores: Context and Texture strict All4 accuracy, Attribute exact counting accuracy and Language Elicitation four-option accuracy; zero-shot, greedy decoding where supported, 8-16 output tokens for API models; 100 generated or edited cases per subset that survived filtering against Gemini 3.5 Flash and human review). Source: arxiv.org. Saturation forecast: Around June 2028. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.6 | 31.3 |
| 2 | Kimi K2.6 (Non-reasoning) | 23.3 |
| 3 | Qwen 3.5 27B (Non-reasoning) | 23 |
| 4 | GPT-5.4 (Non-reasoning) | 18 |
| 5 | Grok 4.3 | 17.8 |
Interactive version: theaggregate.ai/benchmark?slug=sabre-prior · How It Works · Data refreshed daily, snapshot 2026-09-29.