AEPC-QA (Closed-Book): leaderboard
Metric: Accuracy (%) answering from the model's own knowledge on AEPC-QA's 807 four-option French multiple-choice questions from the Quebec financial regulator's (AMF) insurance certification exam-preparation handbooks, answered with a single letter (unparsable or refused answers score as wrong), mean over a stratified 10-fold protocol (seeds 42 to 51); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 51 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | O3 (2025-04-16) | 76.13 | #117 |
| 2 | Gemini 2.5 Pro | 74.65 | #145 |
| 3 | O1 (2024-12-17) | 72.1 | #144 |
| 4 | Claude Opus 4 (20250514) | 66.91 | #148 |
| 5 | GPT-4.5 (Preview) | 66.67 | #239 |
| 6 | Claude Sonnet 4 (20250514) | 66.46 | #211 |
| 7 | Gemini 2.5 Flash | 65.47 | #237 |
| 8 | Claude 3.7 Sonnet (20250219) | 64.98 | #196 |
| 9 | GPT-4.1 | 64.32 | #240 |
| 10 | O4 Mini (2025-04-16) | 64.28 | #173 |
| 11 | Mistral Medium 3 | 62.84 | #483 |
| 12 | Grok 3 | 62.68 | #296 |
| 13 | Magistral Medium | 61.81 | #478 |
| 14 | Llama 3.3 70B Instruct | 61.48 | #520 |
| 15 | Grok 3 Mini | 61.19 | #274 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=aepc-qa-closed-book · How It Works · Data refreshed daily, snapshot 2026-10-11.