Gemini 3 Pro: benchmark results
Google's Pro-tier Gemini 3 model for advanced reasoning and general tasks. Provider: Google. Released 2025-11-18. Access: API.
Unified ELO 1716 ± 1, rank #34 of 1392 rated models, from 295 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ADBench | 83 | Pass@3 (self-reported) | 100 |
| AISE-Bench | 56.06 | F1-LM (%) | 100 |
| BALROG MiniHack (LLM) | 40 | Progress (%) | 100 |
| BALROG NetHack (LLM) | 6.8 | Progress (%) | 100 |
| BacktestBench | 67.41 | Overall Accuracy (OA) (self-reported) | 100 |
| Chatbot Arena (Text - German) | 1522 | Arena Score | 100 |
| Chatbot Arena (Vision - Creative Writing) | 1324 | Arena Score | 100 |
| Chatbot Arena (Vision - Entity Recognition) | 1298 | Arena Score | 100 |
| CreditCardQA | 70.5 | Chain-of-Thought Accuracy (%) | 100 |
| DiffCap-Bench | 83.1 | F1* (self-reported) | 100 |
| Diplomacy: Overall Performance | 60.3 | Score | 100 |
| EmoBench-M | 70.5 | Average Score | 100 |
Interactive version: theaggregate.ai/model?slug=gemini-3-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.