PM-LLM-Benchmark — leaderboard
Process mining benchmark evaluating LLMs across 8 categories: conformance checking, process modeling, querying, hypothesis generation, fairness, optimization, and vision tasks.
Metric: Score. Source: github.com. Status: saturation imminent. 164 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (xHigh) | 40.2 |
| 2 | GPT-5.6 Sol (High) | 39.8 |
| 3 | Kimi K3 | 39.2 |
| 4 | GPT-5.6 Terra (High) | 38.4 |
| 5 | Kimi K2.6 | 37.8 |
| 6 | GPT-5.4 (xHigh) | 37.8 |
| 7 | GPT-5.5 (High) | 37.7 |
| 8 | GPT-5.3 Codex (xHigh) | 37.3 |
| 9 | GPT-5.6 Luna (High) | 37.2 |
| 10 | GPT-5.4 (High) | 37 |
| 11 | GPT-5.5 (xHigh) | 36.9 |
| 12 | GPT-5 Pro | 36.9 |
| 13 | GPT-5.2 (xHigh) | 36.7 |
| 14 | DeepSeek V4 Pro | 36.4 |
| 15 | GPT-5.6 Sol | 36.3 |
Interactive version: theaggregate.ai/benchmark?slug=pm-llm-benchmark · How the rankings work · Data refreshed daily, snapshot 2026-07-22.