Wire
Gradient gate holds safety across six tests
A Unidirectional Safety Gate kept post-fine-tuning attack success at pre-release levels across six model-dataset settings by blocking gradients from harmful samples inside a calibrated protected region. The Gradient Immunity paper tested Qwen3-14B and Llama-3.1-8B across JailbreakBench, HarmBench, and BeaverTails-H; safe pass rates stayed at 100% on the first two datasets but fell to 79% and 69% on the harder BeaverTails setting, exposing the utility trade-off. For open-weight teams, this is a promising release-time defense rather than a solved control: it assumes a protected component and limited downstream fine-tuning, constraints that sharpen the governance premium identified in the open-weight coalition’s expansion.