SmartBench (Context-Dependent) - Anomaly Localization: leaderboard

Metric: Anomaly location score (0-1 times 100): Jaccard similarity between the devices the model blames and the annotated evidence devices, averaged over the anomalous samples of the 2,400 context-dependent samples (1,200 normal, 1,200 anomalous compressed device-event sequences); higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 10 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro36.5#145
2Gemini 3 Pro (Preview)34.7#64
3DeepSeek R126.1#245
4Claude Sonnet 4.525.7#138
5GPT-5 Mini25.2#176
6GPT-525.1#91
7Claude Sonnet 4 (20250514)24.7#211
8Qwen 3 32B18.5#424
9DeepSeek V317#312
10Qwen 3 8B10.5#667

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=smartbench-context-dependent-anomaly-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.