Gemini 2.0 Pro 02-05 preview — benchmark results
Provider: Google. Released 2025-02-05. Access: API.
Unified ELO 1578 ± 63, rank #598 of 1839 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety BBQ | 95.6 | BBQ accuracy (%) | 72.1 |
| HELM AIR-Bench | 68.4 | Refusal Rate (%) | 55.8 |
| HELM Safety Anthropic Red Team | 98.7 | LM Evaluated Safety score (%) | 45.3 |
| HELM Safety XSTest | 95.4 | LM Evaluated Safety score (%) | 44.8 |
| HELM Safety | 90.5 | Mean score (self-reported) | 42.8 |
| HELM Safety HarmBench | 65.3 | LM Evaluated Safety score (%) | 34.9 |
| HELM Safety SimpleSafetyTests | 97.5 | LM Evaluated Safety score (%) | 32.6 |
Interactive version: theaggregate.ai/model?slug=gemini-2-0-pro-02-05-preview · How It Works · Data refreshed daily, snapshot 2026-08-05.