DeEscalWild - Realism (GPT-5.4 Judge): leaderboard
Metric: Realism score (0-100) from a GPT-5.4 judge rating behavioural plausibility, linguistic naturalness and persona adherence of the generated civilian turns, on DeEscalWild's 150 held-out police-civilian interactions transcribed from real social-media footage: at each turn the model, prompted zero-shot with the situation and a civilian persona, receives the gold officer utterance and generates the civilian reply (autoregressive simulation); mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.2 3B Instruct | 35.5 |
| 2 | Qwen 2.5 3B Instruct | 34.8 |
| 3 | granite-3.0-2B-instruct | 31.8 |
| 4 | Gemma 2 2B (IT) | 28.1 |
Interactive version: theaggregate.ai/benchmark?slug=deescalwild-realism-gpt-5-4-judge · How It Works · Data refreshed daily, snapshot 2026-10-07.