SmartBench (Context-Independent): leaderboard

Metric: F1 (%) of anomaly detection (anomalous is the positive class) on the 2,000 context-independent samples (1,000 normal, 1,000 anomalous single-moment device states with environment readings); temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 10 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)79.3#64
2GPT-579#91
3Gemini 2.5 Pro73.5#145
4GPT-5 Mini72.5#176
5DeepSeek R172#245
6Claude Sonnet 4.568.6#138
7Qwen 3 32B64.8#424
8Claude Sonnet 4 (20250514)60.1#211
9DeepSeek V351.3#312
10Qwen 3 8B46.2#667

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=smartbench-context-independent · How It Works · Data refreshed daily, snapshot 2026-10-11.