Position Bias (Lechmazur) — leaderboard

Tests whether LLM judges preserve the same underlying preference when two lightly edited stories are shown in opposite orders. Lower order-flip rates indicate more stable pairwise judgment.

Metric: Order Flip % (lower is better). Source: github.com. Status: saturation imminent. 36 models tracked.

Top models

#ModelScore
1MiMo-V2-Pro19.8
2Seed 2.0 Pro28
3Gemini 3.5 Flash29.8
4Claude Opus 4.6 (Thinking, High)30.2
5DeepSeek V3.230.3
6GLM-5.131.5
7Qwen 3.6 Plus33.8
8Qwen 3.7 Max34.8
9MiniMax-M334.9
10Gemini 3.1 Pro (Preview)35.4
11MiniMax-M2.736.5
12Arcee Trinity Large (Thinking)36.6
13Claude Sonnet 4.6 (Thinking, High)37.4
14Claude Opus 4.6 (Non-reasoning)38.1
15Grok 4.20 0309 (Reasoning)39

Interactive version: theaggregate.ai/benchmark?slug=position-bias-lechmazur · How the rankings work · Data refreshed daily, snapshot 2026-07-22.