StakeBench - Commitment Detection: leaderboard

Metric: Accuracy (%) on G1: does a comment come from a user holding a position in the market (yes or no, chance 50), one-shot prompt, Polymarket and Manifold comments of resolved markets, macro-average over 18 topic and platform splits, greedy decoding; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 15 models tracked.

Top models

#ModelScore
1Qwen 3 32B (Non-reasoning)74.7
2Qwen 3 14B (Non-reasoning)74
3DeepSeek R1 Distill Llama 8B72.1
4Gemma 2 9B71.8
5Gemini 2.5 Flash71.1
6Claude Haiku 4.569.7
7Qwen 3 8B (Non-reasoning)69.2
8GPT-5.568.9
9Qwen 3 30B A3B (Non-reasoning)67
10Llama 3.2 3B Instruct56.7
11Mistral 7B Instruct (v0.3)53.4

Interactive version: theaggregate.ai/benchmark?slug=stakebench-commitment-detection · How It Works · Data refreshed daily, snapshot 2026-10-07.