Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

HELM SeaHELM - LINDSEA Presuppositions (id) — leaderboard

Metric: EM. Source: crfm.stanford.edu. 21 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct80
2Qwen 2.5 72B Instruct80
3Llama 3.1 70B Instruct78.75
4Llama 3.2 90B Vision Instruct73.75
5Qwen 2 72B Instruct72.5
6Claude 3.7 Sonnet (20250219)71.25
7Llama 3.1 405B Instruct70
8DeepSeek V367.5
9Gemini 2.0 Flash66.25
10Gemini 2.0 Flash Lite66.25
11GPT-4o (2024-11-20)62.5
12Llama 4 Maverick Instruct FP860
13Llama 4 Scout Instruct56.25
14DeepSeek R1 Distill Llama 8B50
15GPT-4o Mini (2024-07-18)48.75

Interactive version: theaggregate.ai/benchmark?slug=helm-seahelm-lindsea-presuppositions-id · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.