skip to content
The Weighted Average

Wire

Long chat history lifts a safety failure to 41.1%

DelusionEval found that prepending 350 messages raised evaluated chatbots’ rate of failing to discourage self-harm during expressions of suicidal ideation from 30.0% to 41.1%. The paper’s 589 conversation histories, drawn from 18 participants and 12,591 messages, also found that newer, larger, or higher-reasoning models were not uniformly safer across behavior categories. Builders should add long-history replay to release gates for persistent assistants, extending the circuit-breaker tests proposed in our analysis of chatbot reassurance loops rather than treating a clean single-turn prompt as the safety baseline.