skip to content
The Weighted Average

Wire

PASTABench finds agents hit 40.74% safety timing

PASTABench evaluated 16 language models across 1,139 multi-turn safety trajectories and found the best model achieved optimal intervention timing in only 40.74% of cases. The benchmark paper separates whether an agent should intervene, when it should intervene, and what risk it sees; it also warns that smaller models’ apparent safety can collapse when hazard keywords are neutralized. Teams shipping tool-using agents should read that gap alongside AISI’s 8.2% scope-crossing agent evaluation and test timing, not only final refusal rates.