SABRE-Prior - Language Elicitation: leaderboard

Metric: Multiple-choice accuracy (%; four options for a question that suggests a plausible answer, the correct option being that the requested detail cannot be determined from the image; zero-shot, greedy decoding where supported, 8-16 output tokens for API models; 100 generated or edited cases per subset that survived filtering against Gemini 3.5 Flash and human review). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.658
2Qwen 3.5 27B (Non-reasoning)29
3Grok 4.323
4GPT-5.4 (Non-reasoning)23
5Kimi K2.6 (Non-reasoning)17

Interactive version: theaggregate.ai/benchmark?slug=sabre-prior-language-elicitation · How It Works · Data refreshed daily, snapshot 2026-09-29.