Bio-Saul-Dolphin-Beagle-Breadcrumbs: benchmark results

Provider: Other. Released 2024-04-01. Access: Open.

Unified ELO 1355 ± 15, rank #2440 of 2928 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Medical LLM - PubMedQA75.4Accuracy (%)72.7
Open LLM Leaderboard v1 - TruthfulQA MC246.39MC2 (%) (0-shot)34.3
Open LLM Leaderboard v1 - MMLU50.74Accuracy (%) (5-shot)30.1
Open LLM Leaderboard v1 - HellaSwag75.4Normalized accuracy (%) (10-shot)26.5
Open LLM Leaderboard v1 - WinoGrande69.46Accuracy (%) (5-shot)23.5
Open LLM Leaderboard v1 - ARC Challenge48.63Normalized accuracy (%) (25-shot)23.4
Open Medical LLM - MMLU Professional Medicine47.06Accuracy (%)19.6
Open LLM Leaderboard v1 - GSM8K1.67Accuracy (%) (5-shot)18
Open Medical LLM - MMLU College Biology50.69Accuracy (%)17.6
Open Medical LLM46.73Average Accuracy (%)16.5
Open Medical LLM - MMLU Clinical Knowledge49.43Accuracy (%)16.5
Open Medical LLM - MMLU College Medicine38.73Accuracy (%)16.5

Interactive version: theaggregate.ai/model?slug=bio-saul-dolphin-beagle-breadcrumbs · How It Works · Data refreshed daily, snapshot 2026-09-23.