Edge LLM Leaderboard: Raspberry Pi 5 — leaderboard

Edge LLM Leaderboard: Raspberry Pi 5 evaluates model capability on inference tasks from the linked upstream source with MMLU Accuracy as the primary reported metric.

Metric: MMLU Accuracy (%). Source: huggingface.co. 128 models tracked.

Top models

#ModelScore
1Mistral-7B-Instruct-v0.3 (Q8_0, llama_cpp)43.2
2Mistral-7B-Instruct-v0.3 (Q4_K_M, llama_cpp)42.9
3Mistral-7B-Instruct-v0.3 (Q4_0_4_4, llama_cpp)42.9
4Phi-3-medium-128k-instruct (Q4_K_M, llama_cpp)42.7
5Qwen2.5-14B (Q4_K_M, llama_cpp)42.5
6gemma-2-9b (Q8_0, llama_cpp)42.4
7gemma-2-9b (Q4_0_4_4, llama_cpp)42.4
8Qwen2.5-14B (Q4_0_4_4, llama_cpp)42.1
9Phi-3-medium-128k-instruct (Q4_0_4_4, llama_cpp)42.1
10Mistral-Nemo-Base-2407 (Q4_0_4_4, llama_cpp)41.9

Interactive version: theaggregate.ai/benchmark?slug=edge-llm-leaderboard-raspberry-pi-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.