ParseBench - Semantic Formatting: leaderboard

Metric: Semantic Formatting Score. Source: parsebench.ai. Saturation forecast: Around February 2028. 78 models tracked.

Top models

#ModelScore
1GLM-5.3 Flash79.49
2Claude Opus 5.577.04
3Claude Fable 5.176.52
4Gemini 3.8 Flash (High)76.36
5DeepSeek V4.1 Flash (High)73.48
6Claude Fable 572.62
7Claude Opus 4.871.38
8Claude Opus 4.769.42
9Gemma 4 31B (IT)69.3
10Gemini 3 Flash (High)68.31
11Gemini 3.5 Flash (Minimal)68.09
12GPT-6 Sol (Medium)67.88
13GPT-6 Sol (Non-reasoning)67.32
14Gemini 3.8 Flash (Low)66.99
15Gemma 4 26B A4B (IT)65.1

Interactive version: theaggregate.ai/benchmark?slug=parsebench-semantic-formatting · How It Works · Data refreshed daily, snapshot 2026-09-24.