FireBench - Output Format: leaderboard

Metric: Share of the 1,300 output format samples (%) in which the model performs the expected action: answering in a required format (JSON, XML, delimiters such as boxed answers and adversarial variants) on long-document and reasoning questions, and separating reasoning, code and explanation in coding-agent prompts, verified programmatically; FireBench enterprise and API instruction following; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScoreOverall rank
1GPT-4.186.9#240
2Qwen 3 235B A22B 2507 Instruct65.6#291
3Claude Sonnet 4.564#138
4GPT-5.1 Instant63.3#277
5GPT-5.1 (Medium)62.2#131 (GPT-5.1)
6Llama 4 Maverick Instruct57.9#439
7Kimi K2 Instruct (0905)56.6#247
8GPT-OSS-120B55.3#330
9Kimi K2 (Thinking)54.6#236 (Kimi K2)
10DeepSeek V3.1 Terminus54.3#212
11Qwen 3 235B A22B 2507 (Thinking)39.9#253 (Qwen 3 235B A22B 2507)

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=firebench-output-format · How It Works · Data refreshed daily, snapshot 2026-10-11.