AEPC-QA (RAG): leaderboard
Metric: Accuracy (%) with the authors' retrieval-augmented pipeline (text-embedding-ada-002 retrieval of the top 5 chunks from a 2.6M-token Quebec insurance law and contract corpus, compressed into the prompt) on AEPC-QA's 807 four-option French multiple-choice questions from the Quebec financial regulator's (AMF) insurance certification exam-preparation handbooks, answered with a single letter (unparsable or refused answers score as wrong), mean over a stratified 10-fold protocol (seeds 42 to 51); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | O3 (2025-04-16) | 78.68 | #117 |
| 2 | O1 (2024-12-17) | 75.18 | #144 |
| 3 | Sonar Deep Research | 73.8 | #191 |
| 4 | Claude Opus 4 (20250514) | 73.17 | #148 |
| 5 | Claude 3.7 Sonnet (20250219) | 72.51 | #196 |
| 6 | Claude Sonnet 4 (20250514) | 71.89 | #211 |
| 7 | O4 Mini (2025-04-16) | 70.41 | #173 |
| 8 | GPT-4.5 (Preview) | 69.26 | #239 |
| 9 | GPT-4.1 | 67.94 | #240 |
| 10 | GPT-4.1 Mini | 65.8 | #346 |
| 11 | GPT-4 (0613) | 63.7 | #475 |
| 12 | Grok 3 | 62.92 | #296 |
| 13 | O1 Mini (2024-09-12) | 62.02 | #393 |
| 14 | O3 Mini (2025-01-31) | 60.12 | #232 |
| 15 | Claude 3.5 Haiku (20241022) | 59.01 | #574 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=aepc-qa-rag · How It Works · Data refreshed daily, snapshot 2026-10-11.