Gemini 3.5 Flash (Low): benchmark results

Provider: Google. Released 2026-05-19. Access: API.

Unified ELO 1685 ± 1, rank #184 of 3078 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Chess Puzzles (Epoch AI)45Accuracy (%)93.2
ObviousBench99.31Answer pass³ (%)92.6
Epoch AI - GPQA Diamond88.89Accuracy (%)85.1
OTIS Mock AIME 2024-2588.89Accuracy (%)76.8
Epoch AI - Mystery Game Puzzles28Score76.6
GRIPS-hard65.3Accuracy (%)73.1
FinLifeBench - Financial State - Checkpoint State Accuracy75.4Cell-level state accuracy (%)70
GRIPS91.6Accuracy (%)69.2
FinLifeBench - Financial State - Evidence Recall77.3Recall of gold evidence sessions (%)40
FinLifeBench - Financial State - Granular Change Accuracy41.2GCA@15 (%)40
FinLifeBench - Life-Event History - Event-Anchor F148.8F1 (%; event type with first-establishing session)30
FinLifeBench - Life-Event History - Exact History Match5.3Checkpoints with the exact history (%)30

Interactive version: theaggregate.ai/model?slug=gemini-3-5-flash-low · How It Works · Data refreshed daily, snapshot 2026-09-19.