SAKE - Zero-Shot - Scalable Practices: leaderboard

Metric: Accuracy (%). Source: arxiv.org. 11 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.693
2Claude Opus 4.692.7
3Grok 4.2092.7
4Gemini 3.1 Flash Lite91.9
5GPT-5.491.1
6DeepSeek V3.291.1
7Qwen 3 235B A22B90.4
8Claude Haiku 4.589.6
9Mistral Small 489.6
10Grok 4.1 Fast88.4
11GPT-5.4 Nano88.4

Interactive version: theaggregate.ai/benchmark?slug=sake-zero-shot-scalable-practices · How It Works · Data refreshed daily, snapshot 2026-09-19.