Optimum LLM Perf Leaderboard: leaderboard
Inference measurements of open LLMs by Hugging Face's Optimum-Benchmark: latency, throughput, memory and energy for a 256-token prompt and 64 generated tokens at batch size 1 on a single GPU.
Source: huggingface.co.
Interactive version: theaggregate.ai/benchmark?slug=optimum-llm-perf-leaderboard · How It Works · Data refreshed daily, snapshot 2026-09-05.