Wire
Safety merge cuts harm recognition to 12.9%
Four standard model-merging methods retained 81% to 85% jailbreak refusal but left harm-classification accuracy at no more than 12.9% in a controlled Gemma-3-1B-IT study. The researchers found the two safety task vectors were nearly orthogonal, yet the larger refusal update dominated the merged weights and turned nuanced recognition into broad refusal. The finding reinforces the lesson from OpenAI’s sandbox escape that one visible control is not the whole safety case: teams merging safety fine-tunes should regression-test each behavior separately, not treat a healthy refusal score as proof the rest survived.