Gemini 3 Deep Think — benchmark results
Google Gemini 3 variant using Deep Think mode for harder multi-step reasoning tasks. Provider: Google. Released 2026-02-26. Access: API.
Unified ELO 2186 ± 71, rank #4 of 1776 rated models, from 20 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Google Gemini 3 Deep Think - ARC-AGI-2 | 84.6 | Score (%) | 100 |
| Google Gemini 3 Deep Think - CMT-Benchmark | 50.5 | Pass@8 (%) | 100 |
| Google Gemini 3 Deep Think - Codeforces | 3455 | Elo | 100 |
| Google Gemini 3 Deep Think - GPQA Diamond | 93.8 | Score (%) | 100 |
| Google Gemini 3 Deep Think - Humanity's Last Exam (no tools) | 48.4 | Score (%) | 100 |
| Google Gemini 3 Deep Think - Humanity's Last Exam (search and code) | 53.4 | Score (%) | 100 |
| Google Gemini 3 Deep Think - International Chemistry Olympiad 2025 (theory) | 82.8 | Score (%) | 100 |
| Google Gemini 3 Deep Think - International Math Olympiad 2025 | 81.5 | Score (%) | 100 |
| Google Gemini 3 Deep Think - International Physics Olympiad 2025 (theory) | 87.7 | Score (%) | 100 |
| Google Gemini 3 Deep Think - MMMU-Pro | 81.5 | Score (%) | 100 |
| LiveCodeBench Pro | 3298 | Rating (CF-style) | 100 |
| CritPt | 25.7 | Accuracy (self-reported) | 98.5 |
Interactive version: theaggregate.ai/model?slug=gemini-3-deep-think · How the rankings work · Data refreshed daily, snapshot 2026-07-22.