JANUS - Ordering: leaderboard

Metric: Ordering distortion, signed from -1 to 1: change in one minus the adverse-first rate over expressed favorable and adverse fact pairs, paired distortion: the aspect score of the goal-conditioned response minus that of the neutral response for the same JANUS scenario (six true facts, three favorable and three adverse to an institutional objective, shown in the same order in both conditions), averaged over scenarios; facts are matched to sentences and framing is labelled by a Qwen3-8B judge with thinking disabled; zero-shot at provider or checkpoint default decoding; lower is better. Source: arxiv.org. Saturation forecast: Around 2032. 12 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini0.04
2Qwen 3 32B0.05
3Qwen 3 8B (Non-reasoning)0.05
4GPT-5.40.05
5Qwen 3 14B0.07
6Llama 3.1 8B Instruct0.08
7Qwen 3 32B (Non-reasoning)0.08
8Qwen 3 8B0.09
9DeepSeek V4 Flash (Non-reasoning)0.1
10Qwen 3 14B (Non-reasoning)0.1

Interactive version: theaggregate.ai/benchmark?slug=janus-ordering · How It Works · Data refreshed daily, snapshot 2026-09-29.