P3B3 (pt-BR Prompt): leaderboard
Metric: Variety score (0-100), where 0 is Brazilian and 100 European Portuguese, of the model replies when the first turn asks for Brazilian Portuguese; Gemini-3-Flash judge, Portuguese category-based single-turn prompt, averaged over 203 turns; greedy decoding; lower is better (follows the requested variety). Source: arxiv.org. Saturation forecast: Estimated already saturated. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 4.1 |
| 2 | gemma-4-E4B-it | 4.2 |
| 3 | Ministral-3-14B-Instruct-2512 | 4.2 |
| 4 | Gemma 3 12B (IT) | 4.9 |
| 5 | Llama 3.3 70B Instruct | 6.7 |
| 6 | Gemma 4 31B (IT) | 7.3 |
| 7 | Qwen 3 8B | 7.9 |
| 8 | Apertus-70B-Instruct-2509 | 8.6 |
| 9 | EuroLLM-22B-Instruct-2512 | 9.6 |
| 10 | Olmo 3.1 32B Instruct | 10.5 |
| 11 | Llama 3.1 8B Instruct | 10.6 |
| 12 | Qwen 3.5 27B | 11.9 |
| 13 | Qwen 3.5 9B | 12.2 |
| 14 | Apertus-8B-Instruct-2509 | 14.7 |
| 15 | OLMo 3 7B Instruct | 15.8 |
Interactive version: theaggregate.ai/benchmark?slug=p3b3-pt-br-prompt · How It Works · Data refreshed daily, snapshot 2026-09-29.