ARFBench - Magnitude: leaderboard
Metric: Accuracy (%) on the 76 Magnitude questions (how far the anomaly deviates from the expected counterfactual) of ARFBench, for the time series rendered as chart images (Matplotlib plots, three images for paired questions), with the series description and channel names as text; few-shot multiple-choice prompting with shuffled answer choices, temperature 0.05 (default for models without the setting), reasoning models at medium effort with 2,000 output tokens, blank or invalid answers counted wrong; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 11 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 | 68.4 |
| 2 | GPT-5 | 65.8 |
| 3 | GPT-4o | 61.8 |
| 4 | Claude Opus 4.6 | 57.9 |
| 5 | GPT-5.4 (Medium) | 57.9 |
| 6 | Gemini 3 Pro | 56.7 |
| 7 | Claude Sonnet 4.5 | 53.9 |
Interactive version: theaggregate.ai/benchmark?slug=arfbench-magnitude · How It Works · Data refreshed daily, snapshot 2026-10-07.