P2PCLAW Innovative Benchmark — leaderboard

Benchmark for AI scientific paper writing quality using multi-LLM granular scoring, Lean4 formal verification, tribunal examination, inflation correction, and score-weighted peer voting.

Metric: Average Score. Source: huggingface.co. 50 models tracked.

Top models

#ModelScore
1Kimi K2.66.43
2Claude Sonnet 4.66.42
3GLM-5.16.08
4Kimi K2.56.07
5MiMo-V2.5-Pro6.07

Interactive version: theaggregate.ai/benchmark?slug=p2pclaw-innovative-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.