skip to content
The Weighted Average

AI Economics for Operators

DeepSeek V4 Pro Adds a Time-of-Day Tax

DeepSeek will charge up to $3.96 per million output tokens at peak from August 16; agent teams should schedule around the new price clock.

A laptop computer with a coffee cup on a desk
A laptop computer with a coffee cup on a desk. Photograph by ian dooley

DeepSeek is changing V4 Pro from a flat price into a time-of-day market. From August 16 at 16:00 UTC, the API’s peak output price will be $3.96 per million tokens and its off-peak price $1.98, versus the current $0.87 output rate. The peak price is a derived 4.55× increase—$3.96 ÷ $0.87—so teams using DeepSeek for agents should audit scheduling, queueing, and routing before assuming yesterday’s cost model still describes production.

This is not simply a price hike. It is a new control surface. The day’s Ultrafast analysis makes the complementary point: latency is also a budget variable, not merely a model attribute. Batch jobs can move into off-peak hours; interactive work can route to another model during the peak window; and cache-heavy workloads can remain cheap if their inputs are reused. But the benefit depends on whether an application can tolerate time-zone scheduling and whether the rate change survives the provider’s own next revision.

The cheap tier now has a clock

DeepSeek’s official pricing page says V4 Pro currently costs $0.435 per million cache-miss input tokens, $0.003625 per million cache-hit input tokens, and $0.87 per million output tokens. It says the new peak/off-peak schedule takes effect at 16:00 UTC on August 16, 2026. Peak hours are 01:00 –04:00 UTC and 06:00 –10:00 UTC; all other hours are off-peak. The page says off-peak rates are half the peak rates.

From the new schedule, V4 Pro becomes $1.98 per million output tokens off-peak and $3.96 at peak. The output price therefore rises from $0.87 to $1.98 even in the cheaper window—about 2.28×—and reaches 4.55× the current rate at peak. The input side also changes: cache misses move to $1.32 peak and $0.66 off-peak, while cache hits become $0.044 peak and $0.022 off-peak. The provider is putting a premium on fresh computation while preserving a much lower price for repeated context.

The API remains compatible with OpenAI and Anthropic formats, according to DeepSeek’s quick-start documentation. The page lists V4 Pro and V4 Flash, says the model version is updated behind the same name, and identifies a 1M-token context length with a maximum output of 384K. Those details make migration technically plausible for existing OpenAI-compatible clients. They do not make the economics interchangeable: a client that sends the same request shape may still have a different cache hit rate, concurrency profile, or peak-hour exposure.

The capacity terms matter too. DeepSeek’s rate-limit documentation lists a 500-connection concurrency limit for V4 Pro and 2,500 for V4 Flash, calculated at the account level. It says accounts can request higher concurrency without an additional cost, subject to capacity review. Peak pricing without available concurrency is a bad bargain; a team should test both before promising throughput to an internal customer.

The immediate operator choice is workload segmentation. Long-running research, nightly code analysis, embedding-style repeated context, and evaluation suites can be scheduled outside peak windows. Interactive copilots, support flows, and incident response should either reserve capacity, use a different model, or accept the premium explicitly. A single blended average hides the decision.

The Gemini 3.7 Flash launch makes the comparison more uncomfortable. Google’s introductory price is $3.75 per million output tokens, which is slightly below DeepSeek’s new $3.96 peak output rate, though the models, benchmarks, latency, and context economics differ. DeepSeek still has an off-peak price advantage at $1.98, and its cache-hit input price remains dramatically lower than the listed alternatives. The point is not that one model wins every workload; it is that “cheap model” is no longer a sufficient procurement category.

Route around the expensive hours

A simple monthly calculation exposes the stakes. Suppose a team sends 100 million output tokens each month, evenly across the listed peak and off-peak hours. The exact share depends on traffic and the scheduling policy, but the published windows total 7 hours per day of peak time—3 + 4—leaving 17 hours off-peak. If traffic were uniform, 7/24 of the output would face $3.96 and 17/24 would face $1.98. The blended rate would be $3.96 × 7/24 + $1.98 × 17/24 = $2.56 per million, or about $256 per month for 100 million output tokens.

That $256 is a derived scenario, not a forecast. Real agents cluster demand during business hours, retries can compound it, and queue delays may force work into another window. The arithmetic’s use is diagnostic: a team must know its traffic distribution before it can claim that off-peak billing will save money. If all 100 million tokens land in peak windows, the bill is $396; if they all land off-peak, it is $198. The same token volume now carries a 2× price spread.

The cache rules can dominate the output story for context-heavy agents. At the new peak rate, a cache-hit input token costs $0.044 per million while a cache-miss input token costs $1.32—a 30× gap. Off-peak, the gap is also 30×: $0.022 versus $0.66. That is not permission to manufacture cache hits. It is a reason to preserve stable prompt prefixes, avoid injecting changing content at the front of every request, and measure cache behavior as a production metric.

The OpenAI latency guide makes the broader system point: fewer output tokens, fewer requests, parallel calls, streaming, and not defaulting to an LLM can all improve an application. DeepSeek’s schedule adds another lever—when the call runs. A team that keeps a serial agent loop, emits verbose intermediate text, and runs it at peak hours has three avoidable multipliers in its bill.

What could break the thesis? First, scheduling can increase time-to-result. If moving a batch to off-peak makes a report arrive after the decision it informs, the lower token rate is false economy. Second, an application may not be able to preserve cache prefixes when tools or user context change every turn. Third, provider pricing can change again; the page itself says product prices may vary and DeepSeek reserves the right to adjust them. Fourth, a second model may cost more per token but produce fewer retries or higher accepted-result rates.

The evidence that would change the verdict is a month of observed invoices split by peak/off-peak, cache hit/miss, model, and accepted result. Compare it with a shadow route through another workhorse model. The right metric is cost per accepted task at a specified latency—not the cheapest cell in a pricing table.

That is also the lesson for platform teams building a router. Treat time as an input to routing, alongside model capability, context size, residency, and approval risk. Make the clock visible in telemetry. Set a budget ceiling for peak calls. Allow a fallback when the price or queue exceeds the task’s value. And test the schedule on the same workload that produced the original cost estimate; synthetic uniform traffic will understate the risk.

The earlier price-war analysis argued that token efficiency, not a single benchmark rank, is becoming the operator’s real frontier. DeepSeek’s move extends that idea into calendar efficiency. The best router may now be the one that knows both what to call and when to call it.

Sources