Deep FinResearch Bench (Firm B Template): leaderboard
Metric: Mean overall report-quality score (1-4 scale, 1 poor to 4 excellent) that GPT-5 assigns to AI-written pre-earnings equity research reports on S&P 500 companies (fiscal 2025 Q1 and Q2), each prompted with a rubric of Firm B's professional reports for the same company and quarter; averaged over comprehensiveness, assumption quality, coherence and analytical depth; deep research agents and LLMs with web search; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 Deep Research | 1.57 |
| 2 | Sonar Pro | 1.54 |
| 3 | GPT-4o | 1.4 |
| 4 | GPT-4o Mini | 1.32 |
| 5 | Gemini 2.5 Flash | 1.25 |
Interactive version: theaggregate.ai/benchmark?slug=deep-finresearch-bench-firm-b-template · How It Works · Data refreshed daily, snapshot 2026-10-07.