CodeGemma-7B: benchmark results
Provider: Google. Access: API.
Unified ELO 1387 ± 19, rank #2229 of 2928 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard v1 - GSM8K | 45.49 | Accuracy (%) (5-shot) | 59.5 |
| ShaderMatch | 11.17 | Clone Match Rate (%) | 47.6 |
| EffiBench - NET | 3.02 | Normalized Execution Time | 42.7 |
| BigCode Models Leaderboard | 40.1 | HumanEval Python Pass@1 (%) | 42.4 |
| EvalPlus (HumanEval+ & MBPP+) | 47 | Pass@1 avg (%) | 40.3 |
| Open LLM Leaderboard v1 - MMLU | 56.59 | Accuracy (%) (5-shot) | 39.6 |
| Big Code Memorization - HumanEval pass@1 | 43.29 | HumanEval pass@1 (%) | 35.3 |
| Big Code Memorization - HumanEval pass@50 | 33.16 | HumanEval pass@50 (%) | 35.3 |
| Big Code Memorization - HumanEval-ET pass@1 | 35.37 | HumanEval-ET pass@1 (%) | 35.3 |
| Big Code Memorization - HumanEval-ET pass@50 | 27.79 | HumanEval-ET pass@50 (%) | 35.3 |
| Open LLM Leaderboard v1 - ARC Challenge | 53.92 | Normalized accuracy (%) (25-shot) | 30.5 |
| Open LLM Leaderboard v1 - HellaSwag | 76.73 | Normalized accuracy (%) (10-shot) | 29.4 |
Interactive version: theaggregate.ai/model?slug=codegemma-7b · How It Works · Data refreshed daily, snapshot 2026-09-23.