NEVU - Macro-F1: leaderboard

Metric: Macro-F1 (%) over the directed Level-1 value labels present in the gold annotations on NEVU's sampled test subset, all four unit levels pooled, shared prompt and actor-centric JSON output, long units handled map-reduce; higher is better. Source: arxiv.org. Saturation forecast: Around May 2027. 13 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro39.87#145
2Claude Opus 4.537.25#79
3Claude Sonnet 4.534.73#138
4GPT-5.232#105
5O331.22#121
6GPT-4.131.12#240
7Claude Haiku 4.530.11#271
8Ministral-3-8B-Instruct-251217.17#644
9Qwen 3 8B9.35#667
10Llama 3.1 8B Instruct5.94#1018
11Phi-3.5-mini-instruct5.73#1148

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=nevu-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.