DeEscalWild - ROUGE-L: leaderboard

Metric: ROUGE-L (0-100) of the generated civilian turns against the real civilian turns, on DeEscalWild's 150 held-out police-civilian interactions transcribed from real social-media footage: at each turn the model, prompted zero-shot with the situation and a civilian persona, receives the gold officer utterance and generates the civilian reply (autoregressive simulation); mean of three runs; higher is better. Source: arxiv.org. Saturation forecast: Around December 2027. 5 models tracked.

Top models

#ModelScore
1Llama 3.2 3B Instruct11.3
2Qwen 2.5 3B Instruct11.2
3granite-3.0-2B-instruct10.2
4Gemma 2 2B (IT)7.2

Interactive version: theaggregate.ai/benchmark?slug=deescalwild-rouge-l · How It Works · Data refreshed daily, snapshot 2026-10-07.