The only build to clear every mandatory gate. Official score 89.25/100.
Intentional design · reliability first
Eight models. One identical website brief.
Every model received the same starter, prompt, harness, tool limit and evaluation. Live builds preserve the model output while applying only the disclosed path adaptations required to serve them together.
Category leaders
Different questions have different winners.
These badges describe the strongest result in each measured category. They do not replace the qualification rule or establish a universal model ranking.
Highest blind visual-review result at 19/20. Not qualified because mandatory gates failed.
Highest raw score at 91.75/100 and highest code score at 8.75/10. Not qualified because one mandatory gate failed.
Qualified · 1 build
Cleared the mandatory bar
Qualified results can receive their full official score.
Grok 4.5
Official score 89.25/100 · uncapped
Not qualified · 7 builds
Raw performance, with failed gates disclosed
These entries are ordered by raw score for diagnosis, not presented as an official second-through-eighth ranking.
GLM-5.2
Official score capped at 49/100
F05Form errors and deterministic success
Claude Fable 5
Official score capped at 49/100
F04Dialog focus trap, Escape and focus return
GPT-5.6 Sol
Official score capped at 49/100
F04Dialog focus trap, Escape and focus returnF05Form errors and deterministic successR01-tabletNo horizontal overflow at 768pxA01No serious or critical axe violations
Kimi K2.7 Code
Official score capped at 49/100
F04Dialog focus trap, Escape and focus returnA01No serious or critical axe violations
MiMo-V2.5-Pro
Official score capped at 49/100
F04Dialog focus trap, Escape and focus return
Llama 4 Maverick
Official score 39.25/100 · raw result already below the cap
F02AFiltering, URL state and direct loadingF03Keyboard-operable service detailF04Dialog focus trap, Escape and focus returnF05Form errors and deterministic successA02Page accessibility while modal is open
MiniMax-M3
Official score 31.5/100 · raw result already below the cap
F02AFiltering, URL state and direct loadingF03Keyboard-operable service detailF04Dialog focus trap, Escape and focus returnF05Form errors and deterministic successA01No serious or critical axe violationsA02Page accessibility while modal is open
Visual and code scores were produced by two independent blind AI reviewer agents. Neither received model identities or the other reviewer’s scores. This is a limitation of Round 0 and must remain disclosed when results are cited.