JuICE - Span Detection: leaderboard
Metric: Span-level F1 (times 100, so 0-100) of erroneous-span detection: the judge reads a user query and a full long-form response and marks the spans with cultural or linguistic errors; a predicted span is a true positive when its word-level IoU with a human-validated JuICE error span exceeds 0.15, over 1,050 query-response pairs in seven country-language settings; higher is better. Source: arxiv.org. Saturation forecast: Around July 2028. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 52.25 |
| 2 | GPT-5.5 | 51.12 |
| 3 | Gemini 3 Flash (Preview) | 48.39 |
| 4 | Claude Opus 4.7 | 45.03 |
| 5 | Gemma 4 31B (IT) | 44.5 |
| 6 | Claude Haiku 4.5 | 36.91 |
| 7 | GPT-OSS-120B | 36.57 |
| 8 | Qwen 3 30B A3B 2507 Instruct | 31.28 |
| 9 | GPT-5.4 Mini | 30.79 |
| 10 | Llama 4 Scout Instruct | 28.05 |
Interactive version: theaggregate.ai/benchmark?slug=juice-span-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.