R2ABench (MetaGPT) - Completeness: leaderboard

Metric: Completeness score (1-5): judge rating (GPT-5.4-mini, calibrated against human review) of how fully the view covers key functional requirements, significant requirements, external systems, data stores and constraints (L2), over the 68 R2ABench projects (17 educational and 51 GitHub repositories), generating a PlantUML architecture view from the structured requirements specification with MetaGPT adapted as an analyst, architect, reviewer and refiner agent pipeline under default settings and greedy decoding; rated over the L2-evaluable views, those that pass the L0 render-and-parse gate (the paper samples its judge calibration from the views that pass L0); higher is better. Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1MetaGPT + Claude Sonnet 4.63.39
2MetaGPT + GPT-53.25
3MetaGPT + DeepSeek V3.22.77
4MetaGPT + Qwen3-Coder 480B-A35B2.67

Interactive version: theaggregate.ai/benchmark?slug=r2abench-metagpt-completeness · How It Works · Data refreshed daily, snapshot 2026-10-07.