ARFBench - Macro-F1: leaderboard
Metric: Multiclass macro-F1 (%) over the 750 ARFBench questions, with answers mapped to per-category semantic answer classes, for the time series rendered as chart images (Matplotlib plots, three images for paired questions), with the series description and channel names as text; few-shot multiple-choice prompting with shuffled answer choices, temperature 0.05 (default for models without the setting), reasoning models at medium effort with 2,000 output tokens, blank or invalid answers counted wrong; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 51.9 |
| 2 | GPT-5.4 (Medium) | 51.4 |
| 3 | Gemini 3 Pro | 49.6 |
| 4 | Claude Opus 4.6 | 46.7 |
| 5 | GPT-4.1 | 44 |
| 6 | GPT-4o | 42.4 |
| 7 | Claude Sonnet 4.5 | 37.9 |
Interactive version: theaggregate.ai/benchmark?slug=arfbench-macro-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.