Fireball-Alpaca-Llama3.1.07-8B-Philos-Math-KTO-beta — benchmark results

EpistemeAI's KTO-aligned Llama 3.1 8B Instruct fine-tune from its Fireball Alpaca line, targeting philosophy and math reasoning. Provider: Other. Released 2024-09-12. Access: Open.

Unified ELO 1428 ± 52, rank #1093 of 1776 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval72.74Score86.6
Open Korean LLM Leaderboard490.55Average Score (%)86.5
Open LLM Leaderboard - MATH Level 515.26Score62.4
Open LLM Leaderboard - MMLU-Pro28.26Score52.9
Open LLM Leaderboard - BBH26.9Score42.4
Open LLM Leaderboard - GPQA4.03Score35.2
Open LLM Leaderboard - MuSR4.28Score21

Interactive version: theaggregate.ai/model?slug=fireball-alpaca-llama3-1-07-8b-philos-math-kto-beta · How the rankings work · Data refreshed daily, snapshot 2026-07-22.