AEPC-QA (RAG): leaderboard

Metric: Accuracy (%) with the authors' retrieval-augmented pipeline (text-embedding-ada-002 retrieval of the top 5 chunks from a 2.6M-token Quebec insurance law and contract corpus, compressed into the prompt) on AEPC-QA's 807 four-option French multiple-choice questions from the Quebec financial regulator's (AMF) insurance certification exam-preparation handbooks, answered with a single letter (unparsable or refused answers score as wrong), mean over a stratified 10-fold protocol (seeds 42 to 51); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 51 models tracked.

Top models

#ModelScoreOverall rank
1O3 (2025-04-16)78.68#117
2O1 (2024-12-17)75.18#144
3Sonar Deep Research73.8#191
4Claude Opus 4 (20250514)73.17#148
5Claude 3.7 Sonnet (20250219)72.51#196
6Claude Sonnet 4 (20250514)71.89#211
7O4 Mini (2025-04-16)70.41#173
8GPT-4.5 (Preview)69.26#239
9GPT-4.167.94#240
10GPT-4.1 Mini65.8#346
11GPT-4 (0613)63.7#475
12Grok 362.92#296
13O1 Mini (2024-09-12)62.02#393
14O3 Mini (2025-01-31)60.12#232
15Claude 3.5 Haiku (20241022)59.01#574

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=aepc-qa-rag · How It Works · Data refreshed daily, snapshot 2026-10-11.