Hype Machine
Sonnet 5.5 Halves Rates; Retries Can Spend the Saving
Sonnet 5.5 costs half Opus 5.5 at standard uncached rates. Repeated attempts, review, and migration determine the saving per accepted task.
Claude Sonnet 5.5 offers standard uncached input and output rates at half those of Opus 5.5, according to Anthropic’s pricing table. A reported landing-page comparison puts the bill reduction at 50.5%, but one additional attempt can consume almost the entire saving.
The cheaper executor has to earn the assignment
The immediate opportunity is to move work that has a clear specification and an inspectable finish condition onto a cheaper route. The important qualifier is finished work. A patch that requires another implementation pass, a review that misses the central defect, or a document that needs extensive correction can turn a lower invoice into a more expensive assignment. Token rates give a builder a reason to test; acceptance determines whether the test earns a rollout.
Anthropic’s Sonnet 5.5 announcement describes a model intended to deliver more efficient coding and knowledge work. Its model overview records a September 28 release and the API identifier claude-sonnet-5-5. The concrete decision is therefore how to evaluate an available model against the tasks a team already pays to complete.
Split the assignment into what must be decided and what must be executed. An implementation with established acceptance tests asks the model to reach a defined destination. An open-ended planning request asks it to identify that destination as well. Those workloads deserve separate scorecards. Combining them into one average can hide a cheaper executor’s strengths or a more expensive planner’s value. The distinction should follow the actual assignment, rather than the model’s marketing category.
Our verdict is to pilot Sonnet for bounded implementation, review, and document tasks with explicit acceptance conditions. Keep a higher-capability path for work whose central difficulty is deciding what should be done. That is a proposed routing rule, not a claim that a model family permanently owns either category. A good specification can make a task easier; an expensive model can still misunderstand it; a cheaper model can still find the right solution.
This decision extends the archive’s examination of Sonnet 5.5’s short-prompt caching threshold without repeating it. The earlier piece asks how prompt size changes the tariff. This one asks how specification quality, retries, and review change the price of an accepted result. The Spark token-mix comparison reached the complementary lesson: a lower-looking rate card is meaningful only after the workload’s actual mix enters the calculation.
Teams already running a validated Opus workflow should resist a wholesale replacement. Copy the evaluation harness first. Preserve its inputs, tools, finish conditions, and review standard. Then change the model and effort deliberately. Without that discipline, a faster run may merely have performed less work. The cheaper model earns its place when the finished work survives the same review.
Half the tariff is a starting point
The cleanest comparison comes from Anthropic’s current pricing table. Standard Sonnet 5.5 input and output rates are $2/$10 per million tokens, against Opus 5.5’s $4/$20. For equal uncached input and output volumes, Sonnet’s token bill is half. That arithmetic describes equal token volumes, not equal task outcomes, and it excludes differences in tools, retries, caching, and application infrastructure.
Nate Herk’s landing-page example gives an observable task-level comparison to place beside the rate card. Herk reports Sonnet completing the PerkForm task in 28 minutes at $6.78, versus Opus at 52 minutes and $13.71. We have not rerun that task or audited its complete usage ledger. Treat these as his demonstrated figures, with the limitations of a single task and a video presentation.
Joining the demonstration to the official tariff produces a useful check. The published rates imply a 50% bill reduction at identical uncached token counts; the demonstrated reduction is ($13.71 − $6.78) ÷ $13.71 × 100 = 50.5%, rounded. The observed result is about 0.5 percentage points above the equal-volume tariff prediction. It is consistent with the headline price relationship, but the small difference cannot reveal token efficiency without the underlying input, output, cache, and tool records.
The retry arithmetic is more revealing than the launch slogan. At those demonstrated bills, $13.71 ÷ $6.78 = 2.02 Sonnet attempts would cost as much as one Opus attempt. Two identical Sonnet attempts total $13.56, leaving just $0.15 before they match the Opus bill. That is a conditional budget comparison, not permission to retry indiscriminately. A rerun may consume different tokens, repeat external work, or create review costs the API invoice misses.
Caching also narrows the apparent discount. The pricing table’s cache-read column lists $0.20 per million tokens for both models. The cheaper uncached rates therefore do not halve an entire bill dominated by shared cache reads. Anthropic’s caching guide explains the separate treatment of cache writes and reads. A team should compare those categories independently rather than apply a blanket discount to its monthly invoice.
The appropriate ledger is small enough to build before migration: input, output, cache writes, cache reads, tool fees, retries, elapsed time, and acceptance outcome. Keep human review time beside it rather than hide it inside an enthusiastic productivity claim. Our earlier Opus cost-efficiency analysis made the same distinction between buying tokens and buying useful work. A tariff saving becomes operational value only when it remains after failed runs and corrections.
The integration can lose before the model does
A switch can fail at the request boundary before anyone judges the answer. The Sonnet 5.5 migration guide documents changed thinking settings and unsupported forced tool choice. Code that relied on forcing a named tool needs a different path. A successful isolated assignment does not establish that those request paths still work in production. Test the actual harness, including its error handling, before routing unattended work through the new model.
The relevant replacement is more than a syntax adjustment. Strict tool use constrains tool arguments to a supported schema. It does not prove that invoking the tool was appropriate, that the source data was trustworthy, or that the resulting operation met the user’s goal. Keep application-level checks around side effects and final results. Valid JSON can still encode a mistaken decision, and a well-typed argument can still target the wrong object.
Conversation handling deserves equal attention. Preserved-thinking documentation explains model and account binding and the importance of preserving the history prefix. A router that moves work between models should not assume reasoning blocks remain usable in every direction. Store a legible task state and inspect the provider’s transformations. A silent loss of prior reasoning can change the economics of a route that appeared straightforward in an isolated demonstration.
Effort is another confounder. Anthropic’s effort guide recommends starting well-specified Sonnet 5.5 agentic coding at medium and moving harder work to high. It also says effort levels are recalibrated relative to Sonnet 5. Carrying the old label forward does not establish an equivalent amount of computation. Comparing two models at their defaults is a legitimate product comparison, but it answers a different question from comparing a deliberately tuned workflow.
The Sonnet-specific prompting guide adds an important qualification: at lower effort the model may stop before completing longer agentic tasks, while higher effort or open-ended requests can encourage extra scope. An operator needs both a finish condition and a boundary. More autonomy is useful when it completes the assignment; it is costly when it extends the assignment into work nobody requested.
The counterargument is strong. A premium model that identifies an unstated requirement may prevent a costly wrong implementation. Herk’s comparison includes an open-ended subscriber-planning prompt where Opus’s result was more detailed and specific. That observation supports testing ambiguity separately; it does not establish a universal quality ranking. The archive’s MLPerf RAG accuracy analysis supplies the right caution: passing a published benchmark gate and satisfying a production acceptance gate are different achievements.
Route by the work you can actually check
The practical rollout starts with tasks whose success can be inspected: a requested patch with existing tests, a document transformation with required fields, or a review bounded to stated concerns. Keep the original assignment and reference materials identical. Mark failures before looking at the model name or invoice. Decide which errors are recoverable and which require escalation. This gives the cheaper model a fair trial without lowering the standard to justify the switch.
Include ambiguous work in a separate group. Ask a reviewer to identify whether the model noticed missing assumptions, surfaced conflicting requirements, or sought information that mattered. Do not grade that group solely by word count or how polished the plan looks. Detail can be helpful, but it can also conceal unsupported assumptions. A shorter answer that flags a real blocker may be more useful than a longer plan that invents its way through one.
For each group, compare cost per accepted result and elapsed time to acceptance. Failed attempts belong in the numerator; accepted results belong in the denominator. If a cheaper route creates extra review or repeated external actions, record them. Set an explicit escalation rule based on the type of failure, rather than repeatedly asking the same model to try harder. The landing-page arithmetic shows why the retry margin can disappear quickly even when the initial invoice looks compelling.
Migration work must enter the decision too. Engineering time to repair tool selection, streaming, history replay, and provider-specific behavior has value even when no token meter records it. The what’s-new documentation is a better rollout checklist than a montage of successful outputs. Map each relevant change to a fixture in the harness, then verify an actual end-to-end assignment. A successful model response is only one stage of a successful application.
There are clear reasons to postpone. Keep the current route when no one can define acceptance, when the application cannot distinguish a tool failure from a completed action, or when the evaluation relies on a reviewer choosing whichever answer feels more impressive. Fixing those gaps may save more than switching models. The new option should make the workflow easier to measure, rather than become a reason to excuse its measurement problem.
What would change our verdict? Repeated tests showing Sonnet’s lower invoice accompanied by more rejected work, more correction time, or materially worse handling of ambiguity would narrow its role. Conversely, equal acceptance quality across representative tasks, with lower complete-run cost and dependable completion, would justify expanding it. Those are observable conditions a team can investigate this quarter. The first successful run is a reason to expand the evaluation, not to abandon it.
The same discipline applies to adjacent product changes. Our Gemini 4 Argon access and output analysis separates a larger capacity ceiling from the opportunity to run an authorized production comparison. The OpenAI DevDay workflow assessment asks which operational improvements become measurable inside a real application. A model, an access program, and a development workflow can each create an opportunity, but none supplies the application’s acceptance test.
For Sonnet, the decision is concrete enough to make now: pilot a defined slice of work, preserve the current quality bar, and account for every attempt. Expand the route when the evidence shows a lower cost of accepted results. Keep the premium path where its contribution survives review and justifies its bill. The resulting routing policy should be an explanation of measured behavior, not an allegiance to a model name.
- Pilot well-specified tasks first, keeping the acceptance test identical to the current route and recording every rejected attempt.
- Verify tool selection, thinking-block handling, streaming, and completion against the actual production harness before expanding unattended use.
- Compare the whole invoice and review burden; keep an explicit escalation path when ambiguity or repeated failure erases the saving.