Gemini 3.1 Flash Lite (Non-reasoning): benchmark results

Provider: Google. Released 2026-05-07. Access: API.

Unified ELO 1585 ± 23, rank #708 of 2055 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UAV-DualCog - Landmark Visibility Counting59.9Visibility-count accuracy (%; 1,024 flight videos in which t96.4
UAV-DualCog - Self-Relative Position33.7Answer accuracy (%; 1,024 questions asking where the UAV is 80
UAV-DualCog - Flight Behavior Recognition (Composite)22.8Composite-level behavior accuracy (%; 1,024 first-person UAV65.4
UAV-DualCog - Future Observation Prediction30.1Answer accuracy (%; 1,024 questions asking which view the UA61.8
UAV-DualCog - Landmark-Relative Direction47.6Answer accuracy (%; 1,024 questions asking where a landmark 55.7
Hy-MultiTurn - Action Suppression56.2Importance-weighted score (%) on the 35 action suppression d47.6
Hy-MultiTurn - Constraint Memory34.4Importance-weighted score (%) on the 34 constraint memory di28.6
UAV-DualCog - Landmark-Driven Action Decision35.9Answer accuracy (%; 1,024 questions asking which way the UAV20
UAV-DualCog - Flight Behavior Recognition (Atomic)23.3Atomic-level behavior accuracy (%; the atomic flight actions19.2
LLM Milgram Obedience - Baseline Full-Obedience Rate100Valid baseline sessions ending at 450 V (%; lower is less ob5.3
LLM Milgram Obedience - Baseline Mean Breakoff Voltage450Mean last shock administered in valid baseline sessions (vol5.3
Hy-MultiTurn38.9Importance-weighted score (%): share of the importance-weigh0

Interactive version: theaggregate.ai/model?slug=gemini-3-1-flash-lite-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-29.