SiteOS benchmark Frozen benchmark · July 2026

Category leaders

Different questions have different leaders.

Leader labels describe the frozen raw evidence. They are not a final qualification verdict.

Model results · 4 models

Raw performance, clearly separated from gates

Cards are ordered by aggregate raw score. Each aggregate equally weights Brand, Ops and Launch.

Scores by brief

Brand, operations and launch

Each cell is the median of two repetitions. Raw score leads; the recorded official score remains secondary.

Scores by brief
Open detailed evidence tables

Component comparison

Where model profiles differ

Accessible text comparison of all model components

Run-level detail

All 24 authoritative runs

Failed gates below are the original recorded outputs—not corrected qualification decisions.

Round 1 run-level scores and recorded gate outcomes
Review disclosure

Visual and code scores were produced by isolated, identity-blind AI reviewer agents. Targeted AI adjudication occurred only where the frozen disagreement rules triggered.

Methodology and limits

What these results mean

Scoring

Automated 60, visual 20, code 10, verification 5 and efficiency 5.

Repetition

Two runs per model and brief. Cell scores use medians; suite scores equally weight the three briefs.

Qualification

The original evaluator capped gate-failing runs at 49. The corrected reanalysis still leaves every output not qualified.

Limits

This is a small, AI-reviewed benchmark. Captured inference cost is not total operating cost.