# SiteOS Round 1 methodology

> The procedure below describes the frozen original result. A post-round audit found evaluator requirements not declared in the briefs, and the corrected public reanalysis is published separately in [ROUND1_REANALYSIS.md](./ROUND1_REANALYSIS.md). The original audit remains historical evidence in [R1_GATE_AUDIT.md](./R1_GATE_AUDIT.md).

Round 1 combines existing authoritative automated Playwright scoring with two isolated blinded AI visual-review passes and two isolated blinded AI code-review passes. Review sessions received only immutable candidate-coded packets and were denied identities, provider/model information, run IDs, automated scores, costs, standings, and other reviewer outputs. The frozen native stage string was retained for compatibility; separate provenance records truthfully identify the reviewers as AI.

Targeted isolated AI adjudication was performed only where frozen disagreement rules triggered: visual total difference greater than 4 or brief-component difference greater than 3, and code total difference greater than 3. Adjudicators saw targeted blinded evidence but no prior scores or reasoning. Two-review components use arithmetic means; three-review components use medians.

Each run totals automated /60, visual /20, code /10, verification /5, and efficiency /5. Verification uses observed check/build and matching successful build evidence. Efficiency awards 2 for no intervention, 1 for a clean exit, 1 for at most 20 minutes (0.5 for 20–30), and 1 for cost at most $5 (0.5 up to $10). Mandatory-gate failures cap official scores at 49. Model/brief cells use repetition medians; suite aggregates equal-weight the three briefs. Ranking follows gate pass rate, category coverage, aggregate raw score, consistency, median cost, then elapsed time.

No authoritative run record, automated score, candidate worktree, reviewer bundle, or frozen rule was modified. Paid website-generation execution remained disabled throughout result processing. The corrected reanalysis did not modify the canonical Round 1 dataset or any published builds.
