Wire
Screenshot-to-code agents miss 80% of pattern traps
Five frontier multimodal models followed the repeated pattern in 80.22% of text-style traps, while averaging only 7.89% accuracy on the visible font-size exception. The 1,440-screenshot benchmark also measured 69.78% bias on card-width perturbations; even the best model, Codex-5.3, fell from 68.61% accuracy on cards to 13.89% on text. Builders using the prompt-to-app workflow described in the Codex Sites analysis should add screenshot diffs and localized visual assertions to acceptance tests, because an agent can recognize an anomaly in its reasoning and still emit the pattern-consistent code.