skip to content
The Weighted Average

Wire

Roleplay redirects a robot policy in 44 of 50 tests

Provael’s expanded SmolVLA red-team run found that a roleplay instruction redirected the robot policy outside its safety envelope in 44 of 50 matched trials across ten LIBERO tasks, versus zero redirections in the paired benign controls. The project’s primary results commit reports an 88% rate with a task-clustered 95% confidence interval of 72%–100%, while three visual or scene-injection attacks produced zero redirections in 50 trials apiece; the authors caution that this is one incompletely seeded draw on one policy and suite. The result sharpens the case made by AISI’s agent evaluation crossing scope in 8.2% of runs: physical-AI teams should gate releases on paired adversarial and benign trajectories across tasks, not a handful of prompt-only checks.