Skip to content

The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps
How It Works Guesswork Metrics
The Aggregate
Loading data...

DiffCap-Bench — leaderboard

Metric: F1* (self-reported). Source: benchmarklist.com. 16 models tracked.

Top models

#ModelScore
1GPT-5.275.1
2Kimi K2.5 (Thinking)74.5
3Kimi K2.571.3
4Grok 4.2069.6
5Qwen 3 VL 32B (Thinking)68.7
6Step3 VL 10B66.1
7Qwen 3 VL 8B (Thinking)65.5
8Qwen 3 VL 32B Instruct63.8
9Qwen 3 VL 8B Instruct61.7
10InternVL3.5-8B51.9

Interactive version: theaggregate.ai/benchmark?slug=diffcap-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.