Gemini 2.0 Flash — benchmark results

Google's workhorse Gemini 2.0 model with a 1M context, native tool use and multimodal input/output (December 2024). Provider: Google. Released 2024-12-11. Access: API.

Unified ELO 1565 ± 7, rank #577 of 1776 rated models, from 534 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM MedHELM - SHC-Proxy74.67EM100
HELM SeaHELM - IndicSentiment98.59Macro F1 score100
HELM SeaHELM - TyDiQA81.51SQuAD macro-averaged F1 score100
HELM SeaHELM - XQuAD (Vietnamese)69.86SQuAD macro-averaged F1 score100
LLM Stats (HiddenMath)63Score (%)100
LLM Stats (Natural2Code)92.9Score (%)100
V-STaR - All mAM26.87Score100
V-STaR - Long mAM37.81Score100
V-STaR - Long mLGM56.14Score100
V-STaR - Medium mAM28.99Score100
Open LMM Reasoning - DynaMath - Subject-scientific figure; Avg72Accuracy (%)99
OpenVLM MMMU - Mechanical Engineering63.3Accuracy (%)98.9

Interactive version: theaggregate.ai/model?slug=gemini-2-0-flash · How the rankings work · Data refreshed daily, snapshot 2026-07-22.