FinSheet-Bench: leaderboard

Metric: Accuracy (%) over all 501 questions (16 templates from simple lookups to complex aggregation) on the 24 synthetic private-equity fund portfolio spreadsheets, each serialized into the prompt zero-shot; a question on a file that exceeds the model's context window counts as wrong (effective accuracy); answers checked by exact matching, fuzzy matching and GPT-4o-mini/Gemini 3 Flash adjudication; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)82.4#54
2GPT-5.2 (Thinking)80.4#105 (GPT-5.2)
3Gemini 3 Pro80.2#77
4Claude Opus 4.6 (Thinking)80.2#60 (Claude Opus 4.6)
5Gemini 2.5 Pro78.8#145
6Claude Opus 4.666.7#60
7GPT-5.2 (Non-reasoning)57.7#105 (GPT-5.2)
8GPT-4o54.1#333
9Gemini 2.0 Flash50.1#331
10GPT-3.5 Turbo24.8#849

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=finsheet-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.