MARCA (Portuguese, Orchestrator): leaderboard

Metric: Checklist accuracy (0-1, times 100): share of a question's checklist items (expected entities and attributes) that a GPT-4.1 judge marks satisfied, averaged over MARCA's 52 Portuguese questions and three runs, Orchestrator framework (the model delegates sub-questions to subagents that call web_search and web_scrape); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro88.9
2GPT-5.287.2
3GPT-4.182.5
4Gemini 3 Flash82
5GPT-5 Mini81.9
6Kimi K277.3
7GPT-4.1 Mini73.5
8Qwen 3 235B A22B67.8
9Gemini 2.5 Pro59.8
10Gemini 2.5 Flash51.7
11Qwen 3 30B A3B49.3
12Gemini 2.5 Flash Lite35.4

Interactive version: theaggregate.ai/benchmark?slug=marca-portuguese-orchestrator · How It Works · Data refreshed daily, snapshot 2026-10-07.