Deep Research: benchmark results

Provider: Other.

Unified ELO 1675 ± 27, rank #419 of 2088 rated models, from 15 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SGR-Bench - Row-F141.5Row-level F1 (%) over all tasks: a row counts only when ever90
DeepResearchEval7.28Report quality (0-10; per-task weighted general and task-spe75
SGR-Bench - Constraint-Guided57.27Item-level F1 (%) on the constraint-guided formulation (the 70
DeepResearchEval - Factual Accuracy76.21Verified-correct share (%; checkable statements a GPT-5-mini62.5
LitReviewBench - Claim Support943.3Expert-preference rating (Elo scale, Bradley-Terry aggregati62.5
KDR-Bench - Main Conclusion Alignment63.5Main conclusion alignment (0-100): a 0-10 judge score of how50
LitReviewBench874.1Expert-preference rating (Elo scale, Bradley-Terry aggregati50
LitReviewBench - Literature Coverage882.2Expert-preference rating (Elo scale, Bradley-Terry aggregati50
SGR-Bench54.2Item-F1 (self-reported)50
SGR-Bench - Goal-Oriented51.14Item-level F1 (%) on the goal-oriented formulation of the sa50
KDR-Bench - Key Point Coverage44.1Key point coverage (%): share of the 261 annotated table-gro37.5
LitReviewBench - Paper Structure857.5Expert-preference rating (Elo scale, Bradley-Terry aggregati37.5

Interactive version: theaggregate.ai/model?slug=deep-research · How It Works · Data refreshed daily, snapshot 2026-10-07.