JuICE - Span Detection: leaderboard

Metric: Span-level F1 (times 100, so 0-100) of erroneous-span detection: the judge reads a user query and a full long-form response and marks the spans with cultural or linguistic errors; a predicted span is a true positive when its word-level IoU with a human-validated JuICE error span exceeds 0.15, over 1,050 query-response pairs in seven country-language settings; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 10 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)52.25
2GPT-5.551.12
3Gemini 3 Flash (Preview)48.39
4Claude Opus 4.745.03
5Gemma 4 31B (IT)44.5
6Claude Haiku 4.536.91
7GPT-OSS-120B36.57
8Qwen 3 30B A3B 2507 Instruct31.28
9GPT-5.4 Mini30.79
10Llama 4 Scout Instruct28.05

Interactive version: theaggregate.ai/benchmark?slug=juice-span-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.