Deep FinResearch Bench: leaderboard
Metric: Mean overall report-quality score (1-4 scale, 1 poor to 4 excellent) that GPT-5 assigns to AI-written pre-earnings equity research reports on S&P 500 companies (fiscal 2025 Q1 and Q2), each prompted with a rubric of Firm A's professional reports for the same company and quarter; averaged over comprehensiveness, assumption quality, coherence and analytical depth; deep research agents and LLMs with web search; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | O3 Deep Research | 2.02 |
| 2 | Sonar Pro | 1.7 |
| 3 | Gemini 2.5 Flash | 1.59 |
| 4 | GPT-4o | 1.5 |
| 5 | GPT-4o Mini | 1.42 |
Interactive version: theaggregate.ai/benchmark?slug=deep-finresearch-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.