skip to content
The Weighted Average

Wire

Autonomous CTF agents finish second overall

Autonomous agents finished second overall in a live capture-the-flag competition, outperforming most human teams in an IBM-led study of 41 participants and four autonomous agents. The study’s abstract says human teams increasingly delegated subtasks but were often bottlenecked by prompting and context specification; security teams should treat agent autonomy and human oversight as separate variables to measure, not assume that adding a copilot improves performance. That echoes the archive’s evaluation-governance finding: benchmark rank is useful only when the harness records the controls and workflow around the model.