FIRE (Finance, Qualification Exams): leaderboard

Metric: Accuracy (%) over all exams, weighted by question count, on FIRE's financial-qualification set (about 14,000 multiple-choice questions from 14 professional certification exams); exact option match; each model with its recommended inference settings, else temperature 0.6 and top-p 0.8; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 14 models tracked.

Top models

#ModelScoreOverall rank
1Seed-1.691.78#257
2Gemini 3 Pro91.43#77
3Seed-OSS-36B-Instruct87.55#348
4Qwen 3 235B A22B 2507 (Thinking)87.31#253 (Qwen 3 235B A22B 2507)
5DeepSeek V3.286.1#198
6Kimi K2 (Thinking)85.95#236 (Kimi K2)
7Qwen 3 Max85.88#201
8GPT-5.283#105
9GLM-4.681.7#246
10Claude Sonnet 4.579.01#138
11GPT-OSS-120B68.94#330

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fire-finance-qualification-exams · How It Works · Data refreshed daily, snapshot 2026-10-11.