AI Economics for Operators
Anthropic Cuts Cache Reads 75% and Holds Its Price
Fable 5.1 keeps $10 input and $50 output but drops cache hits to $0.25, cutting agentic bills up to 45% without a headline price cut.
Anthropic shipped Claude Fable 5.1 and Mythos 5.1 on Tuesday without touching the headline price — still $10 per million input tokens and $50 per million output — and instead cut the cache-read rate by 75%, to $0.25 per million. On a long-running agent that re-reads a large repository or document set on every turn, that single line item is worth more than any list-price cut the company has made this year: Anthropic estimates 25% lower cost for typical workloads and up to roughly 45% for highly agentic work, per its Fable 5.1 and Mythos 5.1 launch post.
The mechanic matters more than the discount. Every other Claude model prices a cache hit at 0.1x base input; Fable 5.1 and Mythos 5.1 price it at 0.025x, a multiplier that exists nowhere else in the lineup, as the Claude Platform pricing documentation now states in a footnote. Anthropic did not lower the price of the model. It lowered the price of memory — and in agentic workloads, memory is most of the bill.
The discount hides in the footnote
Do the arithmetic Anthropic did not print. A coding agent working a persistent context typically re-reads the same prompt each turn; a 90% cache-hit rate is unremarkable in that shape. Blend the rates on the pricing page: 0.9 × $0.25 + 0.1 × $10 gives $1.23 per million input tokens on Fable 5.1, against 0.9 × $1.00 + 0.1 × $10 = $1.90 on Fable 5. That is a 35% cut in effective input cost at unchanged list price, and it lands the priciest model in Anthropic’s catalog just $0.28 above Opus 5’s blended $0.95 — a model whose base input rate is half as much.
Anthropic's cache cut makes its priciest model behave like a cheaper one
Blended cost per 1M input tokens at a 90% cache-hit rate, US dollars
The comparison is the point. Anthropic’s premium tier has been failing to sell on price: the archive documented in August that Fable 5 took only 8% of US business model spend while cheaper models absorbed the rest. A 2.5-cent-on-the-dollar cache rate is the cheapest possible answer to that problem, because it discounts precisely the customers Anthropic most wants — agent operators with big, stable contexts — while leaving the sticker intact for everyone comparing spec sheets. Cache hits also stack with the 50% Batch API discount documented on the same pricing page, so a batched, heavily cached pipeline sees the effect twice.
There is a structural asymmetry worth naming. Under the standard 0.1x rule, Anthropic’s prompt-caching documentation notes a 5-minute cache pays for itself after a single read, since the write costs 1.25x. At 0.025x, a Fable 5.1 write at $12.50 per million amortizes across reads that cost almost nothing, which changes the cache-design calculus: longer cache durations and more aggressive breakpoints become rational where they previously were not. Teams that tuned caching against a 0.1x world are now leaving money on the table by default, the same way teams that never instrumented queries per task overpaid on the search layer beneath their agents.
The competitive frame is narrower than it looks. OpenAI’s published API pricing discounts cached input automatically rather than through an explicit write-then-read contract, and its prompt-caching guide applies the discount to prefixes the platform decides to keep. Anthropic’s model is the opposite bargain: the customer pays a premium to write the cache and then controls exactly what stays hot. At 0.025x, that explicit contract finally prices like the automatic one, which removes the last cost argument against building long-lived context into an agent rather than reconstructing it every turn.
Capability arrives with the invoice
The price move would be a gimmick if the model were flat. It is not. Fable 5.1 posts 55.8% on Terminal-Bench 4.0 against Fable 5’s 42.0% and GPT-5.6 Sol’s 37.3%, with Mythos 5.1 at 60.9%, on the benchmark table in Anthropic’s launch post. On Terminal-Bench-Science 0.1 the jump is starker: 52.6% versus 24.7%, more than double, on a suite where the vendor discloses a ±3.5–4.5 point standard error and reproduces public leaderboard numbers within noise. AutomationBench nearly doubles, 31.4% from 17.1%.
Stack capability against the cache cut and the derived figure sharpens. Fable 5’s blended $1.90 bought 42.0 Terminal-Bench points, or $0.045 per benchmark point per million input tokens; Fable 5.1’s $1.23 buys 55.8, or $0.022. Effective cost per unit of agentic coding capability fell 51% in one release, with no change to the number on the pricing page. That is the ratio to bring to a renewal conversation, and it is the same normalization the paper used when Tencent’s Hy4 preview charged 5x GLM for a 2% edge.
Two of Anthropic’s disclosed research results argue the capability jump is not benchmark-shaped alone. Mythos 5.1 wrote custom GPU kernels that sped up seven open-source genomics and protein models by up to 2.5x on an H100, which the company estimates cuts GPU costs 30–60% on analyses that run those models thousands of times — work it says would normally take a team of performance engineers weeks. Fable 5.1 also trained a network that mapped a third of Venus from three-decade-old Magellan radar at two-to-three-kilometer resolution instead of 10 to 20. Neither is a product, but both are the kind of long-horizon task where cache economics decide whether an operator can afford to let an agent run. The same week, Anthropic previewed a Model Hardware Standard for letting Claude operate laboratory equipment directly.
Anthropic also loosened the safeguards that made Fable awkward in security work. The company says its newest cybersecurity safeguards block 60% fewer false positives, and that Fable 5.1 may now be used to discover software vulnerabilities though not to develop exploits for them. Separately, Enterprise Frontier Safeguards will store monitored data in infrastructure the customer controls, with zero-data-retention access available to eligible customers until EFS ships in phases this fall — the concession TechCrunch framed as the release’s least-restrictive change. For regulated buyers who stalled on log retention, the blocker moved.
Where the arithmetic breaks
Start with the cache-hit rate, because the entire derived saving depends on it. Ninety percent is a plausible steady state for a long agent session with a stable prefix; it is a fantasy for a fan-out workload with many short, dissimilar prompts. At a 50% hit rate the blend is $5.13 versus $5.50 — a 7% saving, not 35%. Anthropic’s own “typical workloads” figure of 25% implicitly assumes something well below the agentic case, and the company does not publish the hit-rate distribution behind either number. Any team quoting 45% to its finance partner should first pull its own cache-hit telemetry.
The honesty caveat cuts both ways in the other direction, too. Anthropic commissioned external red-teaming of the Fable 5.1 cybersecurity safeguards from two organizations plus automated testing by Gray Swan, and says it found no critical-severity jailbreak — a claim it has now made across three consecutive releases while separately conceding, in its report on investigating incidents in cybersecurity evaluations, that greater autonomy brings new failure modes. Absence of a found jailbreak is not evidence of absence, and the same loosened false-positive threshold that unblocks legitimate vulnerability research widens the surface that testing has to cover.
Second, output tokens did not move. At $50 per million, Fable 5.1 remains five times its own input rate and twice Opus 5’s output price, and reasoning-heavy agents that emit long traces will find the savings diluted. The launch post notes Fable 5.1 defaults to High effort in Claude Code and Medium elsewhere; effort level is now a cost lever the caching change does not touch. A migration that leaves effort at High may spend the cache savings on thinking tokens before the quarter ends — the mirror image of the capacity repricing the paper traced when Claude Code’s weekly limit “raise” cost 20% more per unit.
Third, the benchmarks are the vendor’s. Anthropic discloses that Fable 5.1 ran with production safeguards enabled and scored zero on tasks where those safeguards intervened, which cuts against its own numbers, and it reproduces public leaderboard scores in-house rather than deferring to them. That is unusually candid, but it is still a self-graded exam. The cyber and alignment picture is likewise mixed: the system card reports Mythos 5.1 as the strongest cyber model the company has released while remaining in the lower risk category of its Frontier Compliance Framework, and TechCrunch quotes the card describing Mythos 5.1 as “a slight regression on overall misaligned behavior compared to Opus 5.” Cheaper context does not make a more capable model easier to supervise, a tension OpenAI priced at roughly 20% of watched inference compute.
Fourth, the discount is revocable. A footnote multiplier is easier to withdraw than a published rate card, and nothing commits Anthropic to 0.025x on the next model. Buyers should treat it as promotional pricing on a strategic workload class, not a permanent structure.
What to do before the next invoice
The cache cut converts a model-selection question into an instrumentation question. If your agents already cache well, Fable 5.1 is a straightforward migration whose savings you can compute today; if they do not, the release is an argument to fix caching before it is an argument to switch models. Evidence that would change this verdict is simple: published hit-rate distributions, or an independent Terminal-Bench reproduction that fails to replicate the 13.8-point gain.
- Platform teams running persistent agents should measure cache-hit rate first, then migrate. The 35% blended saving requires a 90% hit rate; at 50% it is 7%, and the difference is visible in your own usage payloads before you change a single model string.
- Teams on Opus 5 for cost reasons should re-run the comparison. Blended input cost at high cache hit rates is now $1.23 against $0.95, a 29% premium for a model that scores 55.8 to Opus 5’s 52.3 on Terminal-Bench 4.0 — a materially different trade than the 2x base-rate gap implies.
- Anyone quoting the 45% figure should cap effort levels in the same change. Output stays at $50 per million and Fable 5.1 defaults to High effort in Claude Code; unmanaged reasoning length can eat the input savings.
- Regulated buyers should re-open the retention conversation now. Enterprise Frontier Safeguards ships in phases this fall with customer-controlled storage, and zero data retention is available to eligible customers in the interim.
- Security teams should retest their refusal baselines. A 60% reduction in cybersecurity false positives means workflows that previously failed closed may now proceed; re-run your red-team suite before assuming the old guardrail behavior holds.
The wider market is moving the same way in different currencies. OpenAI’s Astra now carries a Critical cyber designation that gates access rather than pricing it, Cognition is raising at a valuation that implies roughly 52 times its run-rate revenue, Palo Alto is selling remediation of a claimed $1 trillion of pre-AI security debt, and Ai2 has shown that a tenth of a benchmark’s questions can reproduce its ranking. In each case the sticker is not where the money is decided.
The larger pattern is that frontier pricing has moved from the sticker to the seams. Vendors now compete on cache multipliers, effort defaults, batch discounts, and retention terms — the fine print that only shows up in an invoice a month later. The price of a frontier model is no longer its price; it is the price of the tokens you re-read.
Sources
- Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1
- Anthropic — Claude Platform model and feature pricing
- Anthropic — prompt caching behavior and multipliers
- Anthropic — Enterprise Frontier Safeguards
- Anthropic — Claude Fable 5.1 and Mythos 5.1 system card
- Anthropic — Frontier Compliance Framework
- Anthropic — Model Hardware Standard research preview
- Anthropic — investigating incidents in cybersecurity evaluations
- OpenAI — API pricing for models and cached input
- OpenAI — prompt caching guide
- TechCrunch — Anthropic’s new Fable release is cheaper, less restrictive