AEPC-QA (Closed-Book): leaderboard

Metric: Accuracy (%) answering from the model's own knowledge on AEPC-QA's 807 four-option French multiple-choice questions from the Quebec financial regulator's (AMF) insurance certification exam-preparation handbooks, answered with a single letter (unparsable or refused answers score as wrong), mean over a stratified 10-fold protocol (seeds 42 to 51); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 51 models tracked.

Top models

#ModelScoreOverall rank
1O3 (2025-04-16)76.13#117
2Gemini 2.5 Pro74.65#145
3O1 (2024-12-17)72.1#144
4Claude Opus 4 (20250514)66.91#148
5GPT-4.5 (Preview)66.67#239
6Claude Sonnet 4 (20250514)66.46#211
7Gemini 2.5 Flash65.47#237
8Claude 3.7 Sonnet (20250219)64.98#196
9GPT-4.164.32#240
10O4 Mini (2025-04-16)64.28#173
11Mistral Medium 362.84#483
12Grok 362.68#296
13Magistral Medium61.81#478
14Llama 3.3 70B Instruct61.48#520
15Grok 3 Mini61.19#274

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=aepc-qa-closed-book · How It Works · Data refreshed daily, snapshot 2026-10-11.