Gemini 1.5 Pro (Sept): benchmark results
September 2024 Gemini 1.5 Pro snapshot, tracked when sources report the dated API model. Provider: Google. Released 2024-09-24. Access: API.
Unified ELO 1546 ± 1, rank #456 of 1392 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RewardBench | 86.78 | Score (%) | 86.1 |
| Deception Effectiveness (Lechmazur) | 0.94 | Deception Score | 76.5 |
| RewardBench Chat Hard | 76.97 | Accuracy (%) | 76.1 |
| LLM Public Goods Game | 24.64 | Avg. Contribution (%) | 73.7 |
| RewardBench Reasoning | 90.22 | Accuracy (%) | 73.4 |
| RewardBench Safety | 85.81 | Accuracy (%) | 66.7 |
| Confabulation Leaderboard (Lechmazur) | 16.83 | Confabulation rate % (lower is better) | 61.9 |
| RewardBench Chat | 94.13 | Accuracy (%) | 48.9 |
| Deception Resistance (Lechmazur) | 0.62 | Vulnerability Score (lower is better) | 47.1 |
| NYT Connections Original | 22.7 | Score (%) | 43.3 |
| AA GPQA Diamond | 58.89 | Accuracy (%) | 33.6 |
| Artificial Analysis Intelligence Index | 4.27 | Intelligence Index | 32.8 |
Interactive version: theaggregate.ai/model?slug=gemini-1-5-pro-sept · How It Works · Data refreshed daily, snapshot 2026-09-05.