MARCA (Portuguese, Basic): leaderboard

Metric: Checklist accuracy (0-1, times 100): share of a question's checklist items (expected entities and attributes) that a GPT-4.1 judge marks satisfied, averaged over MARCA's 52 Portuguese questions and three runs, Basic framework (the model itself calls web_search via the Serper API and web_scrape); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro89.6
2GPT-5.285.5
3GPT-5 Mini85.2
4GPT-4.182.6
5Kimi K279.7
6Gemini 3 Flash73.3
7Qwen 3 235B A22B67.7
8GPT-4.1 Mini63.1
9Gemini 2.5 Pro58.5
10Qwen 3 30B A3B48.4
11Gemini 2.5 Flash47.5
12Gemini 2.5 Flash Lite29.9

Interactive version: theaggregate.ai/benchmark?slug=marca-portuguese-basic · How It Works · Data refreshed daily, snapshot 2026-10-07.