MARCA (English, Basic): leaderboard

Metric: Checklist accuracy (0-1, times 100): share of a question's checklist items (expected entities and attributes) that a GPT-4.1 judge marks satisfied, averaged over MARCA's 52 English questions and three runs, Basic framework (the model itself calls web_search via the Serper API and web_scrape); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScore
1Gemini 3 Pro88.5
2GPT-5.283.7
3GPT-5 Mini83.7
4Gemini 3 Flash82.8
5GPT-4.181.5
6Kimi K276.9
7Qwen 3 235B A22B66.8
8Gemini 2.5 Pro65.8
9GPT-4.1 Mini64
10Gemini 2.5 Flash52.8
11Qwen 3 30B A3B46.9
12Gemini 2.5 Flash Lite26

Interactive version: theaggregate.ai/benchmark?slug=marca-english-basic · How It Works · Data refreshed daily, snapshot 2026-10-07.