R2ABench (OpenHands) - Completeness: leaderboard

Metric: Completeness score (1-5): judge rating (GPT-5.4-mini, calibrated against human review) of how fully the view covers key functional requirements, significant requirements, external systems, data stores and constraints (L2), over the 68 R2ABench projects (17 educational and 51 GitHub repositories), generating a PlantUML architecture view from the structured requirements specification with the OpenHands agent framework under default settings and greedy decoding; rated over the L2-evaluable views, those that pass the L0 render-and-parse gate (the paper samples its judge calibration from the views that pass L0); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.

Top models

#ModelScore
1GPT-53.67
2Claude Sonnet 4.63.6
3Qwen 3 Coder 480B A35B2.97

Interactive version: theaggregate.ai/benchmark?slug=r2abench-openhands-completeness · How It Works · Data refreshed daily, snapshot 2026-10-07.