Deep Research: benchmark results
Provider: Other.
Unified ELO 1675 ± 27, rank #419 of 2088 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SGR-Bench - Row-F1 | 41.5 | Row-level F1 (%) over all tasks: a row counts only when ever | 90 |
| DeepResearchEval | 7.28 | Report quality (0-10; per-task weighted general and task-spe | 75 |
| SGR-Bench - Constraint-Guided | 57.27 | Item-level F1 (%) on the constraint-guided formulation (the | 70 |
| DeepResearchEval - Factual Accuracy | 76.21 | Verified-correct share (%; checkable statements a GPT-5-mini | 62.5 |
| LitReviewBench - Claim Support | 943.3 | Expert-preference rating (Elo scale, Bradley-Terry aggregati | 62.5 |
| KDR-Bench - Main Conclusion Alignment | 63.5 | Main conclusion alignment (0-100): a 0-10 judge score of how | 50 |
| LitReviewBench | 874.1 | Expert-preference rating (Elo scale, Bradley-Terry aggregati | 50 |
| LitReviewBench - Literature Coverage | 882.2 | Expert-preference rating (Elo scale, Bradley-Terry aggregati | 50 |
| SGR-Bench | 54.2 | Item-F1 (self-reported) | 50 |
| SGR-Bench - Goal-Oriented | 51.14 | Item-level F1 (%) on the goal-oriented formulation of the sa | 50 |
| KDR-Bench - Key Point Coverage | 44.1 | Key point coverage (%): share of the 261 annotated table-gro | 37.5 |
| LitReviewBench - Paper Structure | 857.5 | Expert-preference rating (Elo scale, Bradley-Terry aggregati | 37.5 |
Interactive version: theaggregate.ai/model?slug=deep-research · How It Works · Data refreshed daily, snapshot 2026-10-07.