skip to content
The Weighted Average

Models & Open Source

DeepSeek Keeps V4 Pro After Announcing Its Retirement

DeepSeek's live pricing page reverses a planned V4 Pro redirect. Flash output is 69.7% cheaper off-peak, but model identity still needs testing.

A yellow locomotive pulling containers through branching railway tracks
A yellow locomotive pulling containers through branching railway tracks. Photograph by Anirudh

DeepSeek’s current pricing notice says V4 Pro will remain available after September 14 with billing unchanged, reversing the planned redirect still described in its V4.1 Flash announcement. Flash offers a 69.7% lower off-peak output rate, but operators should treat that saving as an optional migration case—not assume a familiar Pro identifier will be replaced on schedule.

Two official pages, two different promises

The contradiction is explicit. The V4.1 Flash launch page says all deepseek-v4-pro requests will route to V4.1 Flash from 04:00 UTC on September 14, at Flash rates, until V4.1 Pro arrives. The pricing page now says customer demand prompted a decision to continue V4 Pro API service after that date, with its billing method unchanged and further notice for subsequent changes. As retrieved for this edition, both statements remain publicly accessible.

The operational reading is that the pricing page supplies the revised stated policy. It is not proof of observed endpoint behavior; this article has not run a live model-identification test. Teams should retain the notice, confirm the intended behavior with their provider contact where available, and verify their own deployments. A release announcement is not a durable substitute for monitoring a production dependency.

Other identifiers have already crossed the boundary in the published documentation. DeepSeek says deepseek-v4-flash and deepseek-v4-flash-vision-exp remain accepted, but their corresponding models have retired and requests are served by V4.1 Flash at Flash prices. That is compatibility of a request name, not preservation of the old model. An unchanged configuration file can therefore coexist with a changed system behind it.

There is a genuine economic reason to evaluate the replacement. The launch’s pricing graphic lists $0.60 per million output tokens off-peak for V4.1 Flash; the live pricing table lists $1.98 for retained V4 Pro. Combining those two records gives (1.98 − 0.60) ÷ 1.98 × 100 = 69.7%, rounded. This is an output-token tariff comparison with both models in the same off-peak window. It is not a measured saving per successful agent task.

The rest of the bill still matters. The pricing table separates cached input, uncached input, and output, and maintains peak/off-peak rates. A migration can change how much reasoning, context, and repair work a task consumes. Keep those quantities in the comparison rather than projecting the output discount onto the entire invoice. A cheaper token is useful only after the system produces an acceptable result with it.

The new architecture is substantive, not merely a renamed endpoint. DeepSeek’s model card describes a multimodal mixture-of-experts model with a causal encoder-decoder design, native image processing, and distinct prefill and decode behavior. Those details explain why an architectural change can support different economics. They do not establish behavioral equivalence with the model a team previously qualified.

Preserve the rollback, not the illusion of a pin

Start the migration audit with the actual identifiers sent by every client. Separate applications already using legacy Flash aliases from those intentionally using Pro. The former need regression checks against a replacement the vendor says is already serving them; the latter need a deliberate choice under the revised retention policy. Treating both groups as one deadline-driven upgrade would obscure their different starting states.

Capacity needs a second check. DeepSeek’s rate-limit documentation lists 2,500 concurrent Flash requests per account versus 500 for Pro, with a request counting until its response completes. Keys under the same account do not create independent pools. The published ceiling is not guaranteed throughput or a latency result; test the traffic shape and handle throttling without assuming that a larger connection allowance makes every job finish faster.

Reasoning settings belong in the test record too. The model card says its instruct benchmarks use maximum reasoning effort and describes different harnesses for different evaluations. Benchmark quality under that configuration does not automatically transfer to a cheaper, shorter, or differently orchestrated production run. Preserve the harness, reasoning configuration, tool permissions, and acceptance rules when comparing versions. Otherwise the model change and the workflow change become impossible to disentangle.

This is a new boundary in the archive’s earlier analysis of DeepSeek’s experimental vision endpoint. The old post treated retirement of that experimental endpoint as evidence that could alter the adoption case. DeepSeek now says the model has retired while its name continues to work. The lesson is not that compatibility aliases are bad; it is that availability of a string is weaker evidence than continuity of the service behind it.

Today’s AWS monitoring lead distinguishes successful execution from successful work. Apply that distinction to this migration. An API response confirms a request was handled, not that the old quality profile survived. Run a small regression set with known outcomes, preserve bad cases, compare total usage, and make the choice reversible where the provider still offers the incumbent. Do not claim rollback to a retired Flash model merely because its alias remains accepted.

The strongest case for switching is that V4.1 Flash improves the customer’s accepted outcomes while reducing total cost. The strongest case for waiting is that Pro already meets a sensitive requirement and the published retention reversal removes the immediate forced-switch rationale. Neither position needs a universal verdict about which model is better. Both need a workload-level result and a clear understanding of the supported endpoint.

Evidence that changes the decision would include a consistent updated migration notice, observed endpoint behavior, stable task-quality results, and invoices matching the expected model rates. Until then, qualify Flash rather than treating it as an invisible upgrade. The 69.7% output discount earns a trial. The disagreement between official notices earns a deployment check before anyone deletes the incumbent path.

Sources