Deep FinResearch Bench: leaderboard

Metric: Mean overall report-quality score (1-4 scale, 1 poor to 4 excellent) that GPT-5 assigns to AI-written pre-earnings equity research reports on S&P 500 companies (fiscal 2025 Q1 and Q2), each prompted with a rubric of Firm A's professional reports for the same company and quarter; averaged over comprehensiveness, assumption quality, coherence and analytical depth; deep research agents and LLMs with web search; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 9 models tracked.

Top models

#ModelScore
1O3 Deep Research2.02
2Sonar Pro1.7
3Gemini 2.5 Flash1.59
4GPT-4o1.5
5GPT-4o Mini1.42

Interactive version: theaggregate.ai/benchmark?slug=deep-finresearch-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.