gemma-2B: benchmark results
Provider: Google. Released 2024-02-21. Access: Open.
Unified ELO 1310 ± 1, rank #1334 of 1392 rated models, from 170 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Finnish NLU - Scandisent FI | 89.49 | Sentiment classification Score (%) | 53.7 |
| Open Japanese LLM - Mbpp Pylint Check | 32.73 | Score (%) | 51.3 |
| Open Japanese LLM - Wiki NER SET F1 | 5.31 | Score (%) | 51.2 |
| EuroEval Dutch NLU - DBRD | 86.49 | Sentiment classification Score (%) | 49.2 |
| AI Energy Score (Text Generation) | 4 | Energy Score (1-5) | 47.5 |
| EuroEval Faroese NLU - FoSent | 23.93 | Sentiment classification Score (%) | 45.6 |
| EuroEval Italian NLU - Sentipolc16 | 47.86 | Sentiment classification Score (%) | 45.5 |
| URIAL-Bench - Math | 3.3 | Judge Score (0-10) | 44.4 |
| EuroEval English NLU - SQuAD | 76.66 | Reading comprehension Score (%) | 38.3 |
| EuroEval Danish NLU - Angry Tweets | 38.91 | Sentiment classification Score (%) | 36.6 |
| EuroEval Faroese NLU - FoQA | 30.06 | Reading comprehension Score (%) | 36.5 |
| EuroEval Spanish NLU - MLQA ES | 53.45 | Reading comprehension Score (%) | 35.9 |
Interactive version: theaggregate.ai/model?slug=gemma-2b · How It Works · Data refreshed daily, snapshot 2026-09-05.