Gemini 3 Flash (Medium): benchmark results
Provider: Google. Released 2025-12-17. Access: API.
Unified ELO 1773 ± 19, rank #182 of 2088 rated models, from 33 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CoopEval - Mediation | 0.87 | Normalized mean payoff (open scale; below 0 when exploited, | 100 |
| PhageBench | 56.33 | Accuracy (%) averaged over the five PhageBench tasks (phage | 100 |
| PhageBench - Contamination Detection | 58.83 | Accuracy (%) on 1,200 two-option items deciding whether a ph | 100 |
| PhageBench - Host Prediction | 62.5 | Accuracy (%) on 1,200 four-option items naming the bacterial | 100 |
| PhageBench - Lifestyle Classification | 57.4 | Accuracy (%) on 1,000 two-option items classifying a phage a | 100 |
| PhageBench - Phage Contig Identification | 70.83 | Accuracy (%) on 1,200 two-option items telling phage contigs | 100 |
| MageBench | 1694 | Rating | 94.7 |
| ParseBench | 75 | Overall Score | 92.5 |
| ParseBench - Tables | 91.01 | GTRM Score (GriTS + TableRecordMatch) | 90.6 |
| ParseBench - Charts | 61.56 | ChartDataPointMatch Score | 84 |
| CresOWLve - Exact Match | 35.08 | Exact-match accuracy (%) after lowercasing and punctuation r | 82.4 |
| CresOWLve - Russian Exact Match | 51.82 | Exact-match accuracy (%) after lowercasing, punctuation remo | 82.4 |
Interactive version: theaggregate.ai/model?slug=gemini-3-flash-medium · How It Works · Data refreshed daily, snapshot 2026-10-07.