Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

MedFrameQA — leaderboard

Metric: Average Accuracy (self-reported). Source: benchmarklist.com. 11 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash54.75
2O350.18
3Claude 3.7 Sonnet49.67
4O4 Mini49.4
5O147.91
6GPT-4 Turbo46.69
7QVQ-72B-Preview46.44
8GPT-4o45.67
9GPT-4o Mini34.55

Interactive version: theaggregate.ai/benchmark?slug=medframeqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.