Scoring
Automated 60, visual 20, code 10, verification 5 and efficiency 5.
Repeated multi-brief evaluation
Round 1 compares raw quality across 24 independently generated websites. The corrected gate reanalysis is final and does not change the qualification outcome.
Category leaders
Leader labels describe the frozen raw evidence. They are not a final qualification verdict.
Model results · 4 models
Cards are ordered by aggregate raw score. Each aggregate equally weights Brand, Ops and Launch.
Scores by brief
Each cell is the median of two repetitions. Raw score leads; the recorded official score remains secondary.
Component comparison
Run-level detail
Failed gates below are the original recorded outputs—not corrected qualification decisions.
Visual and code scores were produced by isolated, identity-blind AI reviewer agents. Targeted AI adjudication occurred only where the frozen disagreement rules triggered.
Methodology and limits
Automated 60, visual 20, code 10, verification 5 and efficiency 5.
Two runs per model and brief. Cell scores use medians; suite scores equally weight the three briefs.
The original evaluator capped gate-failing runs at 49. The corrected reanalysis still leaves every output not qualified.
This is a small, AI-reviewed benchmark. Captured inference cost is not total operating cost.