Bullshit Benchmark — leaderboard

Tests whether models detect nonsense. Questions use fake precision, confident nonsense framing, cross-domain concept stitching, and other deceptive techniques. Rated green/amber/red by multiple judges.

Metric: BS Detection Rate (%). Source: petergpt.github.io. Status: saturated. 162 models tracked.

Top models

#ModelScore
1Claude Opus 4.896.4
2Claude Sonnet 4.694.5
3Claude Opus 4.692.7
4Claude Opus 4.590.9
5Claude Haiku 4.587.3
6Claude Opus 4.780
7Grok 4.20 Multi-Agent67.3
8Claude Sonnet 4.565.5
9Qwen 3.5 397B A17B65.5
10Claude Sonnet 565.5
11nemotron-3-super-120B-a12B65.5
12Grok 4.2061.8
13Grok 4.361.8
14Grok 4.560
15Claude 3.5 Haiku58.2

Interactive version: theaggregate.ai/benchmark?slug=bullshit-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.