# SiteOS Round 0 results

Generated: 2026-07-14T21:57:14.409Z

Harness: OpenCode 1.17.18 pure mode through OpenRouter; 30-minute ceiling; one scored attempt per model.

Scored runs: 8 / 8. Recorded scored-run cost: $6.4498 USD (adapter probes excluded).

> A blank score means the model has not completed the frozen run and blind review. “Not testable” is preferable to silently substituting a route or harness.

| Rank | Company | Model | Status | Gate | Auto /60 | Visual /20 | Code /10 | Verify /5 | Efficiency /5 | Minutes | Cost USD | Total /100 |
|---:|---|---|---|---|---:|---:|---:|---:|---:|---:|---:|---:|
| 1 | xAI / SpaceXAI | Grok 4.5 | completed | pass | 55 | 17.5 | 6.75 | 5 | 5 | 1.9 | 0.3699 | 89.25 |
| 2 | OpenAI | GPT-5.6 Sol | completed | fail | 41 | 19 | 7.25 | 5 | 5 | 8.9 | 1.4915 | 49 |
| 3 | Anthropic | Claude Fable 5 | completed | fail | 51 | 14 | 8.25 | 5 | 5 | 10.3 | 3.9549 | 49 |
| 4 | Moonshot AI | Kimi K2.7 Code | completed | fail | 40 | 17.75 | 8 | 4 | 5 | 9.2 | 0.3175 | 49 |
| 5 | Z.ai / Zhipu AI | GLM-5.2 | completed | fail | 56 | 17 | 8.75 | 5 | 5 | 6.8 | 0.1486 | 49 |
| 6 | Xiaomi MiMo | MiMo-V2.5-Pro | completed | fail | 41 | 10 | 5.5 | 5 | 5 | 9.5 | 0.0920 | 49 |
| 7 | Meta | Llama 4 Maverick | completed | fail | 26 | 8.25 | 0 | 0 | 5 | 0.2 | 0.0034 | 39.25 |
| 8 | MiniMax | MiniMax-M3 | completed | fail | 18 | 3 | 3.5 | 2 | 5 | 5.4 | 0.0720 | 31.5 |

## Cohort disclosure

The original 14 candidates were filtered by a frozen 3/3 adapter gate. These 6 candidates did not enter the ranked benchmark; their outcomes are compatibility/reliability findings, not coding-quality scores.

| Company | Model | Probe result | Disposition |
|---|---|---:|---|
| Google | Gemini 3.1 Pro Preview | 0/3 | route-incompatible |
| DeepSeek | DeepSeek V4 Pro | 2/3 | adapter-unstable |
| Alibaba Cloud / Qwen | Qwen3.7-Max | 2/3 | adapter-unstable |
| Mistral AI | Devstral 2 2512 | 0/3 | provider-error |
| StepFun | Step 3.7 Flash | 2/3 | adapter-unstable |
| NVIDIA | Nemotron 3 Ultra | 0/3 | adapter-incompatible |

Meta Muse Spark 1.1 was replaced before the original freeze because it was absent from the live OpenRouter model catalogue. Meta Llama 4 Maverick is the disclosed Meta entrant; it is not presented as the same model.

## Advancement

Round 0 advancement is decided only after every testable model finishes blind review. The top five gate-eligible models plus any eligible model within three points of fifth advance to three-run Round 1. This breadth screen does not establish a universal coding-model champion.

