Wire
SAGE cuts self-evolving agent regressions to zero
SAGE, a new statistical acceptance gate for self-evolving agents, reduced the naive gate’s regression rate from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA in a study spanning 20 backbone–benchmark settings. The arXiv paper says it compares incumbent and edited skills item by item and accepts changes only when wins beat regressions with a one-sided test; the result is a controlled research evaluation rather than evidence of production reliability. Teams that let agents rewrite prompts, skills, or workflows should add a held-out acceptance gate before persisting edits, alongside AISI’s evidence on agent evaluation scope.