# SiteOS Round 1 results

The corrected Round 1 gate reanalysis removed invalid evaluator assumptions, but the final outcome is unchanged: 0 of 24 outputs qualified and no outcome is undetermined. Original raw scores, descriptive scores, reviews, costs, builds, screenshots, capped recorded scores and historical gate records remain unchanged. See [ROUND1_REANALYSIS.md](./ROUND1_REANALYSIS.md) and [R1_GATE_AUDIT.md](./R1_GATE_AUDIT.md).

Final qualification by model:

- GPT-5.6 SOL: Not qualified — performance score 53.42/100.
- Grok 4.5: Not qualified — performance score 48.50/100.
- Claude Fable 5: Not qualified — performance score 47.50/100.
- Kimi K2.7 Code: Not qualified — performance score 42.75/100.

All 24 authoritative baseline runs were valid. The original evaluator recorded no fully gate-passing run and capped official scores at 49 where applicable. That qualification conclusion is final under corrected brief-faithful gates.

## Frozen recorded standings

| Rank | Model | Official | Raw | Automated | Visual | Code | Gate rate | Coverage | Consistency | Median cost |
|---:|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
| 1 | openai-gpt-5-6-sol | 48.58 | 53.42 | 15.5 | 18.75 | 9.33 | 0% | 0/3 | 96.83 | $1.237 |
| 2 | xai-grok-4-5 | 47.33 | 48.5 | 13.67 | 16.33 | 8.5 | 0% | 0/3 | 95.33 | $0.323 |
| 3 | anthropic-claude-fable-5 | 46.08 | 47.5 | 16.67 | 12 | 9.17 | 0% | 0/3 | 91.33 | $5.319 |
| 4 | moonshot-kimi-k2-7-code | 41.67 | 42.75 | 14.33 | 11.67 | 6.75 | 0% | 0/3 | 90.83 | $0.230 |

## Scores by brief

| Model | Brief | Official | Raw | Automated | Visual | Code | Verification | Efficiency | Spread | Cost range |
|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---|
| openai-gpt-5-6-sol | brand | 49 | 56.25 | 18.5 | 18.75 | 9.5 | 5 | 4.5 | 2.5 | $1.414–$1.623 |
| openai-gpt-5-6-sol | ops | 49 | 54.75 | 17 | 18.75 | 9 | 5 | 5 | 1.5 | $1.134–$1.147 |
| openai-gpt-5-6-sol | launch | 47.75 | 49.25 | 11 | 18.75 | 9.5 | 5 | 5 | 5.5 | $0.892–$1.326 |
| xai-grok-4-5 | brand | 49 | 51 | 18 | 14.25 | 8.75 | 5 | 5 | 4 | $0.318–$0.327 |
| xai-grok-4-5 | ops | 47.75 | 49.25 | 14 | 17.75 | 7.5 | 5 | 5 | 5.5 | $0.382–$0.382 |
| xai-grok-4-5 | launch | 45.25 | 45.25 | 9 | 17 | 9.25 | 5 | 5 | 4.5 | $0.283–$0.299 |
| anthropic-claude-fable-5 | brand | 49 | 49.5 | 21 | 10 | 9 | 5 | 4.5 | 1 | $5.184–$6.317 |
| anthropic-claude-fable-5 | ops | 43.75 | 47.25 | 16 | 12.75 | 9 | 5 | 4.5 | 17.5 | $5.454–$5.846 |
| anthropic-claude-fable-5 | launch | 45.5 | 45.75 | 13 | 13.25 | 9.5 | 5 | 5 | 7.5 | $2.957–$4.554 |
| moonshot-kimi-k2-7-code | brand | 44.75 | 48 | 18 | 12.25 | 7.75 | 5 | 5 | 15 | $0.190–$0.219 |
| moonshot-kimi-k2-7-code | ops | 33.75 | 33.75 | 13 | 6.25 | 4.5 | 5 | 5 | 9.5 | $0.070–$0.431 |
| moonshot-kimi-k2-7-code | launch | 46.5 | 46.5 | 12 | 16.5 | 8 | 5 | 5 | 3 | $0.240–$0.590 |

## Review disagreement and adjudication

Two and only two frozen-rule triggers occurred: `C-E025B423` visual (first-two totals 9 and 16) and `C-1D4345C6` code (first-two totals 9 and 5). Fresh isolated AI adjudicators reviewed only targeted blinded evidence; their component scores were combined by median.

## Failures and limitations

- All 24 authoritative records were valid completed runs; there were no invalid units in the final cohort. Preserved infrastructure/recovery attempts are not additional scored runs.
- Every scored run failed at least one mandatory gate under corrected brief-faithful gates. No output qualified.
- Review is AI-based, not human review. Subjective visual/code judgments remain a limitation despite independent passes and targeted adjudication.
- The cohort has two repetitions per model/brief cell, so consistency estimates are descriptive and small-sample.
- Costs are inference costs already captured in authoritative run records; hosting, review labor, and credit-purchase fees are excluded.
- The original gate audit is preserved as historical evidence.
