Gemini 2.0 Flash — benchmark results
Google's workhorse Gemini 2.0 model with a 1M context, native tool use and multimodal input/output (December 2024). Provider: Google. Released 2024-12-11. Access: API.
Unified ELO 1565 ± 7, rank #577 of 1776 rated models, from 534 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM MedHELM - SHC-Proxy | 74.67 | EM | 100 |
| HELM SeaHELM - IndicSentiment | 98.59 | Macro F1 score | 100 |
| HELM SeaHELM - TyDiQA | 81.51 | SQuAD macro-averaged F1 score | 100 |
| HELM SeaHELM - XQuAD (Vietnamese) | 69.86 | SQuAD macro-averaged F1 score | 100 |
| LLM Stats (HiddenMath) | 63 | Score (%) | 100 |
| LLM Stats (Natural2Code) | 92.9 | Score (%) | 100 |
| V-STaR - All mAM | 26.87 | Score | 100 |
| V-STaR - Long mAM | 37.81 | Score | 100 |
| V-STaR - Long mLGM | 56.14 | Score | 100 |
| V-STaR - Medium mAM | 28.99 | Score | 100 |
| Open LMM Reasoning - DynaMath - Subject-scientific figure; Avg | 72 | Accuracy (%) | 99 |
| OpenVLM MMMU - Mechanical Engineering | 63.3 | Accuracy (%) | 98.9 |
Interactive version: theaggregate.ai/model?slug=gemini-2-0-flash · How the rankings work · Data refreshed daily, snapshot 2026-07-22.