JANUS: leaderboard

Metric: Average distortion, signed from -1.6 to 1.6: mean of the five paired distortion scores (selection, emphasis, ordering, specificity, framing), paired distortion: the aspect score of the goal-conditioned response minus that of the neutral response for the same JANUS scenario (six true facts, three favorable and three adverse to an institutional objective, shown in the same order in both conditions), averaged over scenarios; facts are matched to sentences and framing is labelled by a Qwen3-8B judge with thinking disabled; zero-shot at provider or checkpoint default decoding; lower is better (less goal-driven distortion). Source: arxiv.org. Saturation forecast: Around 2032. 12 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini0.04
2GPT-5.40.05
3Qwen 3 32B0.06
4DeepSeek V4 Flash (Non-reasoning)0.06
5Qwen 3 8B (Non-reasoning)0.07
6Qwen 3 8B0.07
7Qwen 3 32B (Non-reasoning)0.07
8Qwen 3 14B (Non-reasoning)0.07
9Qwen 3 14B0.08
10Llama 3.1 8B Instruct0.1

Interactive version: theaggregate.ai/benchmark?slug=janus · How It Works · Data refreshed daily, snapshot 2026-09-29.