AI Safety & Security
Failure modes, cyber risk, red-teaming, containment, privacy, reliability, and the controls required to deploy AI safely.
Powerful AI systems expand both capability and attack surface. This topic examines cyber offense and defense, prompt and tool abuse, software supply-chain risk, model failure, privacy, red-teaming, containment, reliability, and validation in high-stakes settings.
Coverage prioritizes operational threat models over abstract reassurance. The useful question is how a system can fail, be exploited, or exceed its intended authority—and which evaluations, permissions, sandboxes, and human review practices reduce that risk. Public regulation appears here only when it directly shapes those technical controls.
-
Data Custody Becomes a Frontier Model Feature
-
OpenAI Puts a 20% Compute Tax on Safety Monitoring
-
Google HEIR Makes Encrypted Inference More Buildable
-
Claude Watermarks Make Provenance a Runtime Duty
-
California Puts an AI Cyber Officer in Every Agency
-
Claude Code Auto Mode Needs Hard Denies
-
OpenAI’s Astra Pause Makes Release Risk a Control
-
AISI’s Agent Eval Crossed Scope in 8.2% of Runs
-
Reddit Needs Four Stages for LLM Moderation
-
Three of Four EU AI Transparency Duties Had No Grace Period