COIN-Bench - Correctness: leaderboard

Metric: Correctness (%): share of the questionnaire's inferred consensus statements that a retrieval-augmented verifier (COIN-RAG, TF-IDF plus MiniLM retrieval over the raw comments) finds consistent with the majority view in the discussions; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash90.06#93
2GPT-5.285.66#105
3O380.35#121
4GPT-4o75.75#333
5Qwen 2.5 72B Instruct64.11#436
6GPT-562.65#91
7Qwen 3 30B A3B61.6#488
8Qwen 2.5 14B Instruct60.88#634
9GPT-4.159.05#240
10DeepSeek R1 Distill Qwen 14B58.45#828
11Qwen 3 32B55.26#424
12Qwen 2.5 32B Instruct54.95#491
13DeepSeek R1 Distill Qwen 32B53.9#640
14Claude 3.5 Sonnet53.35#337
15Llama 3.1 8B Instruct52.67#1018

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=coin-bench-correctness · How It Works · Data refreshed daily, snapshot 2026-10-11.