SABRE-Prior - Texture: leaderboard

Metric: Strict All4 accuracy (%; a case counts only when all four yes/no probes on the base and edited image are right: a familiar object given a counterfactual material; zero-shot, greedy decoding where supported, 8-16 output tokens for API models; 100 generated or edited cases per subset that survived filtering against Gemini 3.5 Flash and human review). Source: arxiv.org. Saturation forecast: Around 2028. 6 models tracked.

Top models

#ModelScore
1Kimi K2.6 (Non-reasoning)52
2Qwen 3.5 27B (Non-reasoning)46
3Claude Sonnet 4.640
4Grok 4.328
5GPT-5.4 (Non-reasoning)28

Interactive version: theaggregate.ai/benchmark?slug=sabre-prior-texture · How It Works · Data refreshed daily, snapshot 2026-09-29.