Gemini 2.0 Flash: benchmark results
Google's workhorse Gemini 2.0 model with a 1M context, native tool use and multimodal input/output (December 2024). Provider: Google. Released 2024-12-11. Access: API.
Unified ELO 1584 ± 1, rank #308 of 1392 rated models, from 546 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM MedHELM - SHC-Proxy | 74.67 | EM | 100 |
| HELM SeaHELM - IndicSentiment | 98.59 | Macro F1 score | 100 |
| HELM SeaHELM - TyDiQA | 81.51 | SQuAD macro-averaged F1 score | 100 |
| HELM SeaHELM - XQuAD (Vietnamese) | 69.86 | SQuAD macro-averaged F1 score | 100 |
| LLM Stats (HiddenMath) | 63 | Score (%) | 100 |
| LLM Stats (Natural2Code) | 92.9 | Score (%) | 100 |
| V-STaR - All mAM | 26.87 | Score | 100 |
| V-STaR - Long mAM | 37.81 | Score | 100 |
| V-STaR - Long mLGM | 56.14 | Score | 100 |
| V-STaR - Medium mAM | 28.99 | Score | 100 |
| Open LMM Reasoning - DynaMath - Subject-scientific figure; Avg | 72 | Accuracy (%) | 99 |
| OpenVLM MMMU - Mechanical Engineering | 63.3 | Accuracy (%) | 98.9 |
Interactive version: theaggregate.ai/model?slug=gemini-2-0-flash · How It Works · Data refreshed daily, snapshot 2026-09-05.