FMG-Bench - Perspective-Compare Harness: leaderboard

Metric: Judge-panel score (0-100; faithful disagreement surfaced). Source: github.com. 14 models tracked.

Top models

#ModelScore
1Claude Opus 4.794.57
2GPT-5.494.12
3Kimi K2.693.82
4Qwen 3.6 Plus93.18
5Grok 4.2092.94
6DeepSeek V4 Pro91.93
7GLM-5.191.34
8MiMo-V2.5-Pro90.87
9Gemini 3.1 Pro (Preview)89.75
10Nemotron 3 Super88.31
11MiniMax-M2.784.6
12Mistral Large 384.42
13Seed 2.0 Lite79.16
14Llama 4 Maverick77.09

Interactive version: theaggregate.ai/benchmark?slug=fmg-bench-perspective-compare-harness · How It Works · Data refreshed daily, snapshot 2026-09-19.