CCiV (Form-Aware) - Structural Accuracy: leaderboard

Metric: Structural accuracy (%, 0-100): a poem counts when its line count and the character count of every line exactly match the target template (non-Chinese characters removed) of the standard form; form-aware prompt: the complete structural and tonal template of the standard form is given; 300 prompts (10 historical themes for each of the 30 most frequent Cipai tune patterns), three samples each at temperature 0.7 and top-p 0.95, averaged; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScoreOverall rank
1Qwen Max78.67#367
2GLM-4 Plus70#373
3DeepSeek R166.67#245
4GPT-4o63#333
5QwQ-32B42.67#410
6DeepSeek R1 Distill Qwen 32B28.33#640
7GPT-4o Mini7.33#588

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=cciv-form-aware-structural-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.