Deep FinResearch Bench (Firm B Template): leaderboard

Metric: Mean overall report-quality score (1-4 scale, 1 poor to 4 excellent) that GPT-5 assigns to AI-written pre-earnings equity research reports on S&P 500 companies (fiscal 2025 Q1 and Q2), each prompted with a rubric of Firm B's professional reports for the same company and quarter; averaged over comprehensiveness, assumption quality, coherence and analytical depth; deep research agents and LLMs with web search; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1O3 Deep Research1.57
2Sonar Pro1.54
3GPT-4o1.4
4GPT-4o Mini1.32
5Gemini 2.5 Flash1.25

Interactive version: theaggregate.ai/benchmark?slug=deep-finresearch-bench-firm-b-template · How It Works · Data refreshed daily, snapshot 2026-10-07.