skip to content
The Weighted Average

Wire

PrecogUI cuts GUI-agent failures under perturbation

PrecogUI beat the strongest baseline by 22.4 percentage points under strong UI perturbations, using InterfereBench’s 1,160 task groups across 34 applications and 27,124 annotated screenshots. Its experience pool, next-layout prediction, and closed-loop correction are aimed at the interruptions that make reactive computer-use agents cascade into failure, not at a clean-demo score. Teams evaluating GUI agents should add overlays, layout changes, and recovery checks to acceptance tests; Kimi Desktop’s approval matrix shows the separate supervision path.