Gemini 2.0 Flash (01-21) (Thinking) — benchmark results

January 21 Gemini 2.0 Flash thinking experimental snapshot, kept separate from the generic thinking row. Provider: Google. Released 2025-01-21. Access: API.

Unified ELO 1630 ± 29, rank #412 of 1776 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
U-MATH - Algebra89.44Accuracy (%)98.5
U-MATH83.64Accuracy (%)97
U-MATH - Differential Calculus70.45Accuracy (%)97
U-MATH - Multivariable Calculus83.71Accuracy (%)97
U-MATH - Integral Calculus82.21Accuracy (%)95.5
U-MATH - Sequences & Series88.31Accuracy (%)90.9
U-MATH - Precalculus92.5Accuracy (%)87.9
μ-MATH81.2Accuracy (%)76.5
NYT Connections Original37Score (%)73.3
Generalization V1 (Lechmazur)1.84Avg Rank (lower is better)68.8
Confabulation Leaderboard (Lechmazur)14.85Confabulation rate % (lower is better)65.1
Wolfram LLM Benchmarking Project46Correct Functionality (%)61.8

Interactive version: theaggregate.ai/model?slug=gemini-2-0-flash-01-21-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.