GRACE - Error Categorization: leaderboard

Metric: Error-categorization macro-F1 (%) averaged over the two tracks: identifying the specific four-class faithfulness-error category of each unfaithful step; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)72.08
2Gemini 3 Flash63.36
3Gemma 4 31B58.5
4GPT-5.458.43
5Qwen 3.5 27B54.36
6Qwen 3.5 35B A3B47.15
7GPT-5.4 Mini41.57
8Qwen 3 14B35.17
9Qwen 3 8B27.21
10Llama 3.1 8B26.88

Interactive version: theaggregate.ai/benchmark?slug=grace-error-categorization · How It Works · Data refreshed daily, snapshot 2026-09-29.