QuoteBench: leaderboard

Metric: Task success (%; exact final-state validation over 56 one-shot Bash tasks in 14 operation families, 14 benign controls and 42 hostile payloads; disclosed-boundary generation contract with the reply replayed through one added nested shell parser; each model at its best-observed effort setting, one stored trial). Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.5 (xHigh)100
2Claude Fable 5 (Max)100
3Claude Opus 5 (xHigh)100
4GPT-5.6 Sol (High)98.2
5Claude Opus 4.8 (Max)94.6
6Gemini 3.1 Pro (Preview) (Low)91.1
7Gemini 3.5 Flash (Medium)64.3
8Claude Sonnet 4.6 (High)64.3
9Qwen 3.5 27B (Non-reasoning)30.4
10Qwen 3.5 4B21.4
11Qwen 3.5 9B (Non-reasoning)17.9
12Gemini 3.1 Flash Lite (Preview)14.3

Interactive version: theaggregate.ai/benchmark?slug=quotebench · How It Works · Data refreshed daily, snapshot 2026-09-26.