FIRE (Finance, Qualification Exams) - Life Insurance (CLIQ): leaderboard

Metric: Accuracy (%) on the China Life Insurance questions of FIRE's financial-qualification set (about 14,000 multiple-choice questions from 14 professional certification exams); exact option match; each model with its recommended inference settings, else temperature 0.6 and top-p 0.8; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro89.4#77
2Seed-1.687.65#257
3Qwen 3 Max85.87#201
4DeepSeek V3.284.67#198
5Seed-OSS-36B-Instruct84.67#348
6Qwen 3 235B A22B 2507 (Thinking)84.57#253 (Qwen 3 235B A22B 2507)
7GPT-5.283.95#105
8Kimi K2 (Thinking)82.1#236 (Kimi K2)
9Claude Sonnet 4.581.38#138
10GLM-4.681.38#246
11GPT-OSS-120B70.78#330

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fire-finance-qualification-exams-life-insurance-cliq · How It Works · Data refreshed daily, snapshot 2026-10-11.