P3B3 (pt-BR Prompt): leaderboard

Metric: Variety score (0-100), where 0 is Brazilian and 100 European Portuguese, of the model replies when the first turn asks for Brazilian Portuguese; Gemini-3-Flash judge, Portuguese category-based single-turn prompt, averaged over 203 turns; greedy decoding; lower is better (follows the requested variety). Source: arxiv.org. Saturation forecast: Estimated already saturated. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)4.1
2gemma-4-E4B-it4.2
3Ministral-3-14B-Instruct-25124.2
4Gemma 3 12B (IT)4.9
5Llama 3.3 70B Instruct6.7
6Gemma 4 31B (IT)7.3
7Qwen 3 8B7.9
8Apertus-70B-Instruct-25098.6
9EuroLLM-22B-Instruct-25129.6
10Olmo 3.1 32B Instruct10.5
11Llama 3.1 8B Instruct10.6
12Qwen 3.5 27B11.9
13Qwen 3.5 9B12.2
14Apertus-8B-Instruct-250914.7
15OLMo 3 7B Instruct15.8

Interactive version: theaggregate.ai/benchmark?slug=p3b3-pt-br-prompt · How It Works · Data refreshed daily, snapshot 2026-09-29.