Silicon-Maid-7B: benchmark results

Provider: Other. Access: Open.

Unified ELO 1484 ± 19, rank #1415 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - HellaSwag86.52Normalized accuracy (%) (10-shot)81.5
Open LLM Leaderboard v1 - ARC Challenge68.17Normalized accuracy (%) (25-shot)79.6
Open LLM Leaderboard v1 - TruthfulQA MC261.64MC2 (%) (0-shot)78
Open LLM Leaderboard v1 - GSM8K61.94Accuracy (%) (5-shot)74.9
Open LLM Leaderboard v1 - MMLU64.58Accuracy (%) (5-shot)72.9
Open LLM Leaderboard v1 - WinoGrande79.01Accuracy (%) (5-shot)66.8
Open LLM Leaderboard - IFEval53.68Score64.3
Open LLM Leaderboard - MuSR41.88Score59.4
Open LLM Leaderboard - GPQA29.03Score45
Open LLM Leaderboard - MMLU-Pro30.83Score39.3
Open LLM Leaderboard - MATH Level 56.5Score35.3
Open LLM Leaderboard - BBH41.28Score26.1

Interactive version: theaggregate.ai/model?slug=silicon-maid-7b · How It Works · Data refreshed daily, snapshot 2026-09-23.