Agentic Engineering
Sonnet 5.5 Cuts Short-Prompt Reuse Cost by 32.5%
Sonnet 5.5 makes 512-token prefixes cacheable. One write and one hit cut that input cost 32.5%, but migration and output bills still need testing.
Teams running short, repeated agent prompts should test Anthropic’s newly released Sonnet 5.5 for its smaller cache threshold, not merely its claimed up to 30% lower cost per task. For a newly eligible prefix, one cache write followed by one hit reduces that portion of input spending by 32.5%—a narrower, reproducible saving that excludes output, tools and migration work.
The discount begins where the old cache stopped
Sonnet 5.5 keeps Sonnet 5’s standard prices: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens, according to the launch. There is no general tariff cut between those generations. Anthropic attributes its broader task-cost claim to using fewer tokens, alongside output generation more than 30% faster. Those are vendor test results, not a bill reduction this publication has independently measured.
A less theatrical change creates a concrete purchasing opportunity. The migration guide lowers the minimum cacheable prompt from 1,024 tokens to 512. Prefixes between 512 and 1,023 tokens can now qualify where they could not on Sonnet 5. That matters to applications with compact, reusable instructions and tool definitions, provided the actual cached prefix reaches the threshold. It does not make every short message cacheable.
Combine that eligibility change with Anthropic’s price schedule: $2 for fresh input, $2.50 for a five-minute cache write and $0.20 for a read, per million tokens. Two uncached passes over the same prefix cost at a combined rate of $4. A write followed by a hit costs $2.70. The proportional saving is (4 − 2.70)/4 = 32.5%. Prefix length cancels from the ratio, so it holds throughout the newly eligible range.
This is a conditional tariff calculation, not an assumed traffic pattern. It requires an actual second request that reads the same cached prefix during its valid lifetime. It excludes any changing suffix and all generated output. If no read follows, the cache write is more expensive than ordinary input. The model launch creates eligibility; the application still has to supply reuse. Finance should not multiply the entire API bill by this discount.
The prompt-caching documentation requires identical cached segments and exposes separate creation, read and uncached-input usage fields. It also says a below-threshold cache request can be processed without caching and without an error. A successful response therefore does not establish that a discount occurred. Check the usage record rather than relying on the presence of a cache setting in the outgoing request.
This extends our earlier analysis of Fable’s cache-read economics without borrowing its workload assumptions. The new Sonnet opportunity is specifically the shorter eligibility threshold. Large prefixes already qualifying on Sonnet 5 do not acquire a new read-price discount simply because the model name changes. For those users, the economic case rests on accepted-task quality, token consumption and latency instead.
Preserve the prefix, then measure the work
The first engineering move is an inventory, not a prompt rewrite. Find repeated prefixes that currently miss the old threshold, and count them under the intended model. Anthropic’s token-counting endpoint accepts structured message inputs and is free within its rate limits. Its documentation also lists unsupported inputs, including most server tools. For those paths, use actual response usage rather than pretending a local character count reproduces the provider’s accounting.
Put stable content before request-specific content and mark the stable boundary deliberately. The cache guide warns that a breakpoint placed on a changing final block will not magically reuse stable material behind it. Timestamps, per-request context and incoming messages belong outside the prefix being compared when they differ. This is a proposed implementation discipline, not a claim that every existing harness has that defect. Inspect the actual serialized request before changing it.
Do not pad a prompt reflexively to manufacture eligibility. Additional instructions can change behavior as well as token volume. First ask whether reusable definitions or policies already belong in the request and have simply been arranged poorly. Then compare accepted results before and after the change. A lower cached-input bill is not a useful saving if a newly crowded prompt makes the agent miss an instruction or require more review.
Output remains a separate cost center. The model overview lists a one-million-token context and a 128,000-token maximum output; neither is a recommendation to use the full allowance. The migration guide says thinking consumes the output budget and is billed as output even when its text is not displayed. A short visible answer can therefore coexist with a substantial generation bill. Record usage, not the apparent brevity of the final message.
Anthropic’s effort guide recommends medium for well-specified agentic work, despite the Claude API default being high. It says effort levels are recalibrated relative to Sonnet 5. Keeping an old label is not a controlled experiment in equal computation. Set effort explicitly, hold the task and acceptance rule fixed, and compare the least costly setting that still completes the work correctly.
The launch makes a strong case for doing that experiment. Anthropic reports substantial coding and knowledge-work gains, while explicitly warning that benchmarks capture only one aspect of capability and that Opus remains stronger on complex, open-ended work. Treat the launch as evidence to form a shortlist, not authority to replace every route. The relevant question is whether this workload benefits from Sonnet’s balance, not whether one headline score makes the entire premium tier unnecessary.
A cache hit cannot rescue a broken contract
Migration can fail before quality enters the conversation. The guide says Sonnet 5.5 rejects thinking: {"type": "disabled"}; applications wanting no up-front thinking must use between_tools, at high effort or below. It also rejects forced tool selection. These are documented request-contract changes, not vague reasons to distrust a new model. A production switch that only edits the model identifier can turn a well-behaved workload into failed requests.
Strict tool use guarantees schema-conforming tool inputs, but that is not a guarantee that the model will call a tool. The migration guide recommends automatic selection with strict tools where supported, and explicitly notes that strict structured outputs are not available for Sonnet 5.5 on Bedrock. A workflow that depended on forced invocation must test the downstream action, not merely remove the parameter that causes an error.
Conversation preservation creates another boundary. Anthropic’s preserved-thinking documentation ties Sonnet 5.5 thinking blocks to their originating account or a linked account. Replaying them through an unrelated account can succeed after those blocks are dropped. The HTTP status remains healthy, but the model no longer receives the same reasoning context. Account switching and saved-session restoration therefore belong in the acceptance test.
The same documentation recommends append-only histories and explains prefix-binding enforcement for newer accounts. This has an economic consequence as well as a compatibility consequence: edits that invalidate reasoning state can also disrupt cache reuse. Test session resumption, tool-list changes and instruction updates on the account and platform that will serve real traffic. A clean fresh conversation is necessary evidence, but it is not sufficient evidence for a persistent agent.
Refusals must also be counted as outcomes rather than transport errors. The refusal guide documents HTTP 200 responses with a refusal stop reason and distinguishes per-attempt billing. Partial refused output is incomplete work. If an authorized fallback serves the task, record the serving model and all billed attempts. Do not credit Sonnet with an accepted result delivered elsewhere, or hide failed work inside an average calculated only from successful requests.
The strongest counterargument is that this is too much machinery for a modest input saving. Sometimes it is. For output-heavy work or rarely repeated short prompts, the newly available cache may barely move total cost. Our Opus 5.5 migration analysis made the same distinction between a token bargain and a working integration. Sonnet should not inherit a lengthy engineering project solely because a small denominator produces an attractive percentage.
Promote a measured route, not a launch default
Start with an existing, frequently repeated task whose result can be checked outside the generating model. Keep the instructions, tools, environment and reviewer standard fixed while evaluating the new model. Then test the cache arrangement as a separate change. Otherwise, a better bill could come from fewer completed actions, different prompts or an altered acceptance threshold. The comparison should make those explanations visible instead of collapsing them into a launch verdict.
Use the first run to establish the write and the next eligible run to establish the read. Preserve the usage records and the accepted result together. If the read field stays empty, investigate prefix identity, timing and eligibility before forecasting savings. If caching works but output grows, calculate the complete task cost rather than celebrating the isolated input line. This article supplies a tariff boundary; production evidence supplies the business case.
Timing deserves particular care. The default cache lifetime is five minutes, while the cache documentation offers a more expensive one-hour write option. The 32.5% calculation uses the five-minute write tariff, not the longer-duration tariff. Applications with intermittent human replies should not import it unchanged. Measure the actual gap between repeated requests, then choose the duration and compare the relevant rates. A technically correct cache hit can still be an economically poor configuration.
There are alternatives to optimizing this meter. Today’s Ninja brief examines a fixed-capacity contract whose nominal employee-hours still need utilization evidence. Today’s Eleven v4 brief separates a launch promotion from the rate a voice product must survive later. Each offer changes a different part of the invoice. Buyers should compare durable accepted work rather than letting a discount, a flat subscription or a benchmark choose the architecture for them.
Evidence that would strengthen the Sonnet switch is straightforward: repeated cache hits in the newly eligible band, preserved acceptance quality, acceptable latency and lower complete-task spending after engineering costs. Evidence that would weaken it is scarce reuse, extra correction, a required tool contract that no longer works, or state loss during realistic resumes. None of those tests requires believing that the vendor’s typical workload resembles yours.
- Agent-platform owners: trial Sonnet 5.5 on compact, repeated prefixes between 512 and 1,023 tokens. Verify a real cache write and hit before applying the 32.5% input-only saving to a forecast.
- Integration owners: budget for the documented thinking, tool-selection and conversation-state changes. Keep an incumbent route available until resumed sessions and required actions pass the same acceptance checks.
- Finance and service owners: compare total input, output, tools, fallback and review cost per accepted task. Expand only when that ledger improves; keep expensive effort settings for work that proves it needs them.
Sonnet 5.5 merits a targeted evaluation now, especially where compact repeated instructions previously fell below the caching floor. It does not establish a universal discount. A smaller cache threshold changes the bill only when the application preserves something worth reading twice.
Sources
- Anthropic — Sonnet 5.5 launch, unchanged tariffs and vendor performance claims
- Anthropic — migration changes and the lower cache threshold
- Anthropic — standard input, output and cache pricing
- Anthropic — cache lifetimes, exact matching and usage accounting
- Anthropic — token counting and unsupported request inputs
- Anthropic — Sonnet context and output limits
- Anthropic — effort recommendations and recalibration
- Anthropic — strict tool-schema guarantees
- Anthropic — account-bound reasoning and conversation preservation
- Anthropic — refusal outcomes and fallback billing