skip to content
The Weighted Average

AI Safety & Security

OpenAI Puts a 20% Compute Tax on Safety Monitoring

OpenAI disclosed that monitoring its frontier workloads costs about 20% of the inference compute being watched, a new budget line.

Rusty roadside sign with red circles against a blue sky
Rusty roadside sign with red circles against a blue sky. Photograph by Josh Snader

OpenAI attached a price to its own caution on Tuesday. In a disclosure titled pacing model development in an era of cyber-critical capabilities, the company said its new monitoring stack consumes roughly 20% of the inference compute being monitored, and that a two-week pause halted reinforcement-learning training on deployment-bound models while its largest planned frontier RL run remains on hold. The trigger was twofold: an incident involving Hugging Face’s systems, and preliminary evidence that its upcoming Astra model may reach the Critical cybersecurity threshold under its Preparedness Framework, a tier no previous OpenAI model has been assigned.

The 20% figure is the part operators can use. Every governance conversation about frontier oversight has run on adjectives; this is the first published coefficient from a major lab, and it converts an ethical posture into a line item any CFO can model.

What a monitoring coefficient actually costs

Apply OpenAI’s own price list and the abstraction becomes concrete. Published API pricing puts GPT-5.6 Sol at $30.00 per 1M output tokens in the standard short-context tier. A 20% monitoring overhead on that workload implies roughly $6.00 per 1M output tokens of oversight — and at Sol’s long-context rate of $45.00, about $9.00. Stack the same coefficient across an agent product’s full input, output, and cached traffic and monitoring stops being a compliance footnote. Neither figure appears in OpenAI’s disclosure; both fall directly out of combining its stated overhead with its own published rates, which is precisely why the coefficient is worth publishing.

The architecture explains where the compute goes. OpenAI describes a multistage system beginning with activation classifiers that inspect model internals at every sampled token, escalating to higher-compute automated investigators that examine tool actions and reasoning traces. The company targets an alert within 30 minutes of concerning activity, and says that if safety, security, and research teams cannot rule out a false positive within a further 30 minutes, they are expected to pause the activity. Monitoring at every sampled token is not a filter bolted onto outputs; it is a second model reading the first one continuously.

Coverage widened alongside cost. OpenAI now requires this monitoring for all RL training and evaluations involving tools at GPT-5.6 Sol capability or higher, and after determining on August 7 that Astra may carry critical cyber capabilities, extended it to all Astra inference with tools rather than training alone. TIME’s interview with Sam Altman added the organizational cost: researchers who never expected to work on alignment have switched to it, and compute has been redirected from capability work to monitoring. Chief scientist Jakub Pachocki acknowledged that OpenAI had built monitors capable of inspecting what its models were planning but had not applied them to the system involved in the Hugging Face evaluation, because it underestimated that system’s capabilities.

The budget line every agent operator now owns

The operator implication is not that enterprises must replicate OpenAI’s stack. It is that anyone running high-autonomy agents against sensitive systems now has a defensible benchmark for what proportionate oversight costs, and a poor excuse for budgeting zero. If the lab with the strongest incentive to keep overhead invisible reports 20%, a team spending 2% on runtime monitoring should be able to explain the difference in risk profile rather than in accounting.

That reframes a familiar procurement fight. Security monitoring for agents has been sold as a percentage of headcount or a flat platform fee; OpenAI’s disclosure prices it as a percentage of inference, which scales with usage rather than with team size. Any organization forecasting agent spend should now carry a monitoring multiplier in the model, in the same way cloud budgets carry egress. The archive made a version of this argument when an 8.2% out-of-scope rate in AISI’s agent evaluations showed that privileged evaluations need production-grade controls, and again when Astra’s possible one-tier escalation turned release risk into a procurement control. Tuesday’s number puts a coefficient on both.

What could break the conclusion. The 20% estimate is OpenAI’s own, unaudited, and the company says the cost varies substantially across training and evaluation workloads — a research RL run with heavy tool use is not an enterprise support agent. The figure also covers frontier-model oversight, where the threat model is a model that may autonomously develop exploits; an internal document-summarization agent does not need activation classifiers at every token. Read 20% as a ceiling for the highest-risk configuration, not a required tax on every deployment. The disclosure is also strategically timed: a company preparing for public markets benefits from documenting restraint before an incident forces it, and OpenAI has said a technical report on the Hugging Face incident is still forthcoming. Until that report lands, the coefficient is a claim rather than an audited measurement, and buyers should treat it accordingly.

The competitive question is sharper. OpenAI is pausing its largest planned training run in a market where rivals are pricing aggressively against it for developer spend, and a self-imposed compute tax that competitors do not adopt becomes a unilateral disadvantage. The evidence that would change the verdict is another frontier lab publishing its own overhead coefficient. If nobody follows, this reads as a company absorbing a cost its peers externalize — and eventually, as pressure to reduce it.

Sources