PM-LLM-Benchmark: leaderboard

Process mining benchmark evaluating LLMs on text-only tasks: contextual understanding of process data, conformance checking and anomaly detection, generating and modifying declarative and procedural process models, process querying, hypothesis generation, identifying unfairness, and process optimization. The diagram-interpretation category is defined upstream but is not part of the scored leaderboard, which reads event logs and process models as text.

Metric: Score. Source: github.com. Status: years away from saturation. 171 models tracked.

Top models

#ModelScore
1GPT-5.6 Sol (xHigh)40.2
2GPT-5.6 Sol (High)39.8
3Kimi K339.2
4GPT-5.6 Terra (High)38.4
5Kimi K2.637.8
6GPT-5.4 (xHigh)37.8
7GPT-5.5 (High)37.7
8DeepSeek V4 Flash (0731)37.7
9GPT-5.5 Pro37.7
10GPT-5.3 Codex (xHigh)37.3
11GPT-5.6 Luna (High)37.2
12GPT-5.4 (High)37
13GPT-5.5 (xHigh)36.9
14GPT-5 Pro36.9
15GPT-5.2 (xHigh)36.8

Interactive version: theaggregate.ai/benchmark?slug=pm-llm-benchmark · How It Works · Data refreshed daily, snapshot 2026-09-05.