StakeBench - Revealed Side: leaderboard

Metric: Strict accuracy (%) on G2: which side (YES or NO) a positioned commenter holds, abstaining counted as wrong, Polymarket and Manifold comments of resolved markets, macro-average over 18 topic and platform splits, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 15 models tracked.

Top models

#ModelScore
1Llama 3.2 3B Instruct53.4
2GPT-5.544.8
3Qwen 3 32B (Non-reasoning)43
4Qwen 3 14B (Non-reasoning)42.3
5Gemma 2 9B39.7
6Gemini 2.5 Flash38.4
7Claude Haiku 4.538.1
8DeepSeek R1 Distill Llama 8B34.5
9Qwen 3 30B A3B (Non-reasoning)34.4
10Mistral 7B Instruct (v0.3)33.5
11Qwen 3 8B (Non-reasoning)32.5

Interactive version: theaggregate.ai/benchmark?slug=stakebench-revealed-side · How It Works · Data refreshed daily, snapshot 2026-10-07.