Agentic Engineering
Codex's Five-Hour Cap Returns to $20 Plus Seats
OpenAI reinstated the rolling five-hour limit on Codex and ChatGPT Work for Plus, keeping Pro exempt. One reporter's four-day run implies a 24x monthly subsidy.
OpenAI has restored the rolling five-hour usage limit on Codex and ChatGPT Work for Plus subscribers, after weeks in which only a weekly cap applied. Engineering lead Thibault Sottiaux said on X, in a post 9to5Mac quotes in full, that the short window “allows us to smoothen the load on our compute” while keeping weekly usage generous, and that Pro $100 and Pro $200 subscriptions keep the cap disabled “for the upcoming months.” The change is a compute-allocation decision dressed as a UX fix, and it lands hardest on exactly the tier most likely to be running an agent all afternoon.
The subsidy the cap is defending
Start with what a Plus seat actually buys. OpenAI’s Codex pricing documentation puts Plus at 10–100 local GPT-5.6 Sol messages per five-hour window, rising to 25–200 on Terra and 250–2,000 on Luna; Pro 5x multiplies those by five and Pro 20x by twenty. Local messages and cloud chats share the same window, and ChatGPT Work draws on the identical budget — the docs state plainly that Work “uses the same pricing, credits, and usage limits as Codex.”
Now price the usage. TechCrunch’s reporter, testing ChatGPT Work on a $20-a-month plan, burned more than 80 million tokens in four days at a cost of about $65 by the model’s own accounting — “a subsidy of more than 3x the subscription price for four days of casual use alone.” Extend that run rate: $65 over four days is $16.25 a day, or roughly $487 over a 30-day month against a $20 seat. That is a 24× gap, and it is the number Sottiaux’s “smoothen the load” is really about. A five-hour window is the cheapest available instrument for capping the tail of that distribution without touching the headline price.
The arithmetic checks against list rates. At OpenAI’s published API pricing — GPT-5.6 Sol at $4 per million input tokens and $20 per million output, with cached input at $0.40 — 80 million tokens costs $65 only if the overwhelming majority are cached or cheap input. Run the same volume as uncached input on Sol and the bill is $320; as output, $1,600. Agentic work is input-heavy and cache-friendly, which is the only reason a $20 seat can absorb any of it at all, and also the reason a single badly scoped long-horizon task can blow through a week’s allowance in an afternoon.
Who should change what, this week
The exemption is the tell. Pro keeps unlimited five-hour access “for the upcoming months,” so OpenAI is not short of compute in general — it is rationing the tier where, in Sottiaux’s words, users “are relatively casual and new” and accidentally consume a week’s usage. If your team runs long agent sessions on Plus seats, the effective ceiling just dropped from a weekly budget you could spend in one sitting to a rolling window you cannot.
Three moves follow, and none of them requires waiting for a vendor announcement. First, route by model, not by habit: Luna’s 250–2,000 messages per window is 25× Sol’s floor, and the docs recommend switching down for routine tasks. Most agent turns do not need frontier reasoning. Second, measure per-window burn before buying seats. The /status command in the Codex CLI and the usage dashboard expose remaining capacity; a team that knows its per-engineer window consumption can price the Pro 5x upgrade against buying credits, rather than guessing. Third, keep an API-key path warm. The same documentation notes that any user can run local chats with an API key at standard rates, which converts a hard cap into a variable cost — worth the accounting friction for a release week.
The counterpoint is real: metered caps are a legitimate answer to a genuinely scarce resource, and a rolling window is gentler than a weekly cliff a user hits on Tuesday. Sottiaux’s second reason — that Plus users blow their whole week accidentally and find the experience confusing — is a defensible product argument, not just a cost story. The pricing page also marks Sol’s current rates as promotional at least through 21 November 2026, so the underlying unit economics are moving the customer’s way even as the guardrails tighten. A cap that arrives alongside cheaper tokens is a different animal from one that arrives alone; the question is which of the two OpenAI extends past the promotional window.
What would change the verdict is disclosure. OpenAI has never published the token allowance behind a Plus seat, only message ranges spanning an order of magnitude, which makes capacity planning guesswork. Until it does, treat subscription agents as best-effort capacity and anything with a deadline as an API workload. That is the same conclusion our analysis of JetBrains’ approach to controlling AI spend reached from the vendor side, and it rhymes with today’s lead on Hugging Face’s $13 billion neutrality premium: the convenience layer is priced by whoever owns it, and the price changes when their compute bill does. Teams that already sized their agent budgets against OpenAI’s Ultrafast latency tiers should rerun the numbers with a five-hour ceiling in place.