Cinder-Phi-2-V1-F16-gguf: benchmark results

Provider: Other. Access: Open.

Unified ELO 1408 ± 20, rank #2065 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - GSM8K47.23Accuracy (%) (5-shot)60.7
Open LLM Leaderboard v1 - ARC Challenge58.28Normalized accuracy (%) (25-shot)39.4
Open LLM Leaderboard v1 - WinoGrande74.66Accuracy (%) (5-shot)38
Open LLM Leaderboard - GPQA28.19Score36.8
Open LLM Leaderboard v1 - MMLU54.46Accuracy (%) (5-shot)34
Open LLM Leaderboard - BBH43.97Score31.6
Open LLM Leaderboard v1 - TruthfulQA MC244.5MC2 (%) (0-shot)27.4
Open LLM Leaderboard v1 - HellaSwag74.04Normalized accuracy (%) (10-shot)24.3
Open LLM Leaderboard - MMLU-Pro21.61Score22.7
Open LLM Leaderboard - IFEval23.57Score17.2
Open LLM Leaderboard - MATH Level 52.42Score15.9
Open LLM Leaderboard - MuSR34.35Score10.1

Interactive version: theaggregate.ai/model?slug=cinder-phi-2-v1-f16-gguf · How It Works · Data refreshed daily, snapshot 2026-09-23.