ARFBench - Macro-F1: leaderboard

Metric: Multiclass macro-F1 (%) over the 750 ARFBench questions, with answers mapped to per-category semantic answer classes, for the time series rendered as chart images (Matplotlib plots, three images for paired questions), with the series description and channel names as text; few-shot multiple-choice prompting with shuffled answer choices, temperature 0.05 (default for models without the setting), reasoning models at medium effort with 2,000 output tokens, blank or invalid answers counted wrong; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 11 models tracked.

Top models

#ModelScore
1GPT-551.9
2GPT-5.4 (Medium)51.4
3Gemini 3 Pro49.6
4Claude Opus 4.646.7
5GPT-4.144
6GPT-4o42.4
7Claude Sonnet 4.537.9

Interactive version: theaggregate.ai/benchmark?slug=arfbench-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.