FIRE (Finance, Qualification Exams) - Actuary (CAA): leaderboard

Metric: Accuracy (%) on the China Actuary (CAA) questions of FIRE's financial-qualification set (about 14,000 multiple-choice questions from 14 professional certification exams); exact option match; each model with its recommended inference settings, else temperature 0.6 and top-p 0.8; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek V3.295.35#198
2GPT-5.294.88#105
3Kimi K2 (Thinking)94.88#236 (Kimi K2)
4Qwen 3 235B A22B 2507 (Thinking)94.42#253 (Qwen 3 235B A22B 2507)
5Seed-1.693.95#257
6GPT-OSS-120B91.16#330
7Gemini 3 Pro91.16#77
8GLM-4.689.77#246
9Seed-OSS-36B-Instruct88.83#348
10Claude Sonnet 4.569.77#138
11Qwen 3 Max65.89#201

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fire-finance-qualification-exams-actuary-caa · How It Works · Data refreshed daily, snapshot 2026-10-11.