JANUS - Framing: leaderboard

Metric: Framing distortion, signed from -2 to 2: change in the mean framing label of matched sentence and fact pairs (1 goal-favoring, 0 neutral, -1 goal-disfavoring), paired distortion: the aspect score of the goal-conditioned response minus that of the neutral response for the same JANUS scenario (six true facts, three favorable and three adverse to an institutional objective, shown in the same order in both conditions), averaged over scenarios; facts are matched to sentences and framing is labelled by a Qwen3-8B judge with thinking disabled; zero-shot at provider or checkpoint default decoding; lower is better. Source: arxiv.org. Saturation forecast: Around 2033. 12 models tracked.

Top models

#ModelScore
1GPT-5.4 Mini0.15
2DeepSeek V4 Flash (Non-reasoning)0.17
3Qwen 3 32B (Non-reasoning)0.18
4Qwen 3 8B (Non-reasoning)0.18
5Qwen 3 14B (Non-reasoning)0.19
6Qwen 3 8B0.2
7GPT-5.40.2
8Qwen 3 32B0.21
9Qwen 3 14B0.22
10Llama 3.1 8B Instruct0.23

Interactive version: theaggregate.ai/benchmark?slug=janus-framing · How It Works · Data refreshed daily, snapshot 2026-09-29.