skip to content
The Weighted Average

Consumer & Creative AI

Eleven v4's Full-Request Discount Is Just $0.58

Eleven v4's launch rate saves $0.58 on a full 10,000-character request. Voice teams should budget at standard rates and test real streaming latency.

Studio microphone in a shock mount beside a round pop filter
Studio microphone in a shock mount beside a round pop filter. Photograph by Amin Asbaghipour

ElevenLabs launched Eleven v4 and its low-latency Turbo variant, but voice teams should separate the model decision from the temporary API discount. At the published request ceiling, that promotion saves $0.58 per v4 generation; it is a cheap opportunity to evaluate quality, not a durable price on which to rebuild a product’s margin.

Price the script after the promotion

The original calculation combines the model documentation’s 10,000-character limit for Eleven v4 with the API page’s $0.022 promotional and $0.08 listed standard rates per 1,000 characters. A full request contains ten billing units. Multiplying gives $0.22 at the launch rate and $0.80 at the standard rate, a difference of $0.58. These are character charges before tax and other products, not the cost of a finished call or approved recording.

Eleven v4's launch discount saves $0.58 per full request

USD per 10,000 characters, before tax. Launch offer listed through October 12.

Launch offerStandard$0$0.2$0.4$0.6$0.8$1$0.8$0.22$0.58 less
Launch offerStandard$0$0.4$0.8$0.8$0.22$0.58 less
ElevenLabs API pricing and model documentation · September 29, 2026

The pricing page labels the offer as running until October 12. Its listed standard rate is 0.08/0.022 = 3.64x the promotional rate, rounded. That is the useful stress test for a product budget. It is not a prediction that the vendor cannot extend or change the promotion; it is the relationship between the two rates displayed when retrieved. A production decision should remain viable without assuming that a temporary offer lasts.

The absolute saving also puts evaluation effort in perspective. A cheap generation is valuable when a team needs to audition many voices or scripts. It does not establish that the resulting material needs less editing, that pronunciation is acceptable, or that a rejected take will not need regeneration. Keep the charge for every attempt and the time spent approving output together. The cheapest character rate can lose to a more expensive model that reliably produces an acceptable take.

ElevenLabs describes better emotional delivery, speaker identity and multilingual consistency, with support for more than 90 languages across the new family. Those are vendor claims worth testing with native-language reviewers and actual production text. They are not evidence that every language, accent and voice combination performs equally well. Choose samples that contain the names, amounts and instructions the application cannot afford to pronounce ambiguously.

The ordinary v4 and Turbo routes also serve different jobs. The model docs present v4 for expressive narration and dialogue, and Turbo for real-time speech. Do not transfer the full-request calculation above into a claim about every Turbo session: Turbo has its own character rate, and a live conversation’s bill depends on what it generates. The comparison here deliberately holds one model and one text size fixed rather than averaging unlike products into a single promotional saving.

A fast voice is not yet a fast conversation

Latency needs equally careful boundaries. The launch reports roughly 150ms median time to first speech for v4 Turbo, with network latency measured and removed. It separately describes approximately 100ms median inference latency. These are different measurements. Neither represents the full delay from a caller finishing a sentence to hearing a useful answer through an application, network and audio device.

The integration documentation makes one source of extra delay concrete. ElevenLabs’ dialogue WebSocket guide describes buffering until roughly 40 characters and eight words arrive, with a flush control for shorter buffers. If an upstream language model emits a short acknowledgment and then pauses, the speech path may wait even though the synthesizer itself is fast. Test streaming behavior with the text cadence the application actually produces.

The same guide says a connection closes after 20 seconds without a client message and documents a keep-alive mechanism. It also says v4 Turbo accepts exactly one registered voice per connection, while ordinary v4 supports more. Those are implementation constraints, not defects. They become expensive only when a team assumes that a successful single utterance proves a long-running, multi-speaker product has migrated correctly.

A sensible trial preserves the whole interaction: incoming speech, transcription, language-model response, text buffering, synthesized chunks and playback. Measure the time until the user hears a relevant response, not merely the provider’s first audio chunk. Also inspect interruptions, short replies, idle gaps and reconnections. This is a recommended evaluation protocol; the retrieved launch does not supply an end-to-end benchmark for the buyer’s complete stack.

Our Muse analysis separated audience scale from useful product activity. The equivalent distinction in creative audio is generated characters versus accepted delivery. A launch can improve both reach and quality without settling the economics of a specific production workflow. Today’s Sonnet lead likewise isolates a narrow input saving from total accepted-task cost, rather than allowing one attractive rate to stand in for the whole invoice.

The strongest case for switching now is a workflow where expression or multilingual consistency is already causing rework. If v4 reduces rejected takes at the standard tariff, the promotion makes qualification cheaper without determining the verdict. For live agents, the case is stronger only when Turbo improves the complete conversation under realistic streaming conditions. A polished demo or a network-excluded median cannot answer that on its own.

Keep the rollout reversible. Preserve the incumbent voice route, compare approved outputs under the same review rubric, and budget using the listed standard rate unless a contract says otherwise. Evidence that would change the verdict is fewer rejected takes, lower editing effort or better measured conversational responsiveness sufficient to cover the durable bill. Switch those proven workloads; delay a blanket migration. A $0.58 discount buys an experiment. It does not buy the evidence the experiment must produce.

Sources