skip to content
The Weighted Average

AI Economics for Operators

Meta's $0.18 Transcription Omits Word Timestamps

Muse Voice Transcribe undercuts a comparable Deepgram feature basket by $0.288 per hour, but missing word timestamps can block a migration.

A tabletop conference microphone with empty chairs behind it
A tabletop conference microphone with empty chairs behind it. Photograph by Daniel Emale

Meeting-intelligence teams have a cheaper candidate in Meta’s newly highlighted Muse Voice Transcribe API, priced at $0.18 per audio hour. Against Deepgram’s current multilingual streaming rate with speaker diarization added, the derived gap is $0.288 per hour—but Meta’s lack of word-level timestamps can make that saving irrelevant to an existing product.

The speaker label is included; the word clock is not

Meta’s model page quotes $3 per 1,000 minutes, or $0.18 per hour, and advertises live attribution for more than 20 speakers. This is a speech-to-text offer, not a complete voice agent: the output supplies a transcript and speaker information, while any downstream reasoning or spoken response requires its own implementation and budget.

The comparator needs matching billing components, not rounded headline prices. Deepgram’s pricing page lists Nova-3 Multilingual streaming at $0.0058 per minute on Pay As You Go, marked as a limited-time promotional rate. Streaming speaker diarization adds $0.0020 per minute. Add those rates and multiply by 60: $0.468 per hour. Subtract Meta’s $0.18 and the difference is $0.288, or approximately 61.5% of the Deepgram feature basket.

This calculation joins the two vendors’ current price pages. It excludes other add-ons, negotiated plans, and any work needed to correct or transform results. It also does not say that the two APIs produce equivalent outputs. The point is to price transcription plus speaker attribution honestly before checking whether those outputs satisfy the same product requirement.

The decisive qualification appears in Meta’s speech-to-text documentation: timestamps are turn-level, not word-level. The guide also says confidence scores, sound-event detection, emotion detection, and transcript reformatting are unavailable. A product that highlights individual words against playback cannot assume the timestamp interface it already uses will survive a model swap.

That missing feature is not a general verdict against Muse. Meeting summaries may need speaker-attributed turns without needing a timestamp for every word. A transcript-review interface may have the opposite requirement. The acquisition decision therefore starts with the application’s output contract, not with a price ranking. Replacing a backend while silently dropping a user-visible capability is a product change, not a cost optimization.

Language coverage needs similar precision. The API guide lists 25 supported languages with code-switching. Meta’s research account says training covered more than 70 languages, with 25 extensively verified. Those statements describe different scopes. A team serving a particular language should test it, rather than turning the larger training count into a promise of equal production support.

Choose the mode before choosing the vendor

Speaker attribution and conversational turn-taking are not the same job. Meta exposes push-to-talk, endpointing, and diarization modes, fixed when the session begins. Its documentation says diarization is not tuned for low-latency voice commands and directs that use case toward endpointing. A buyer cannot infer that enabling the richest-looking mode improves every voice interaction.

The developer walkthrough explains that diarization covers non-overlapping speech. That makes interruptions and people speaking over each other important pilot cases. Do not assume a large advertised speaker count establishes accurate separation when those speakers talk simultaneously. Speaker capacity, overlap handling, and correct attribution are distinct properties even when a marketing sentence places them close together.

Our earlier MAI-Transcribe analysis separates cheap processing from the cost of a usable transcript. This release adds another version of that distinction: a lower raw price can still require output adaptation. Today’s Qwen media-agent budget analysis takes the argument downstream, where retrieval, reasoning, and actions can cost more than the audio input itself. Neither component rate is an end-to-end product margin.

For a meeting-notes application, begin with approved recordings that reflect the real vocabulary, languages, microphone conditions, and speaker transitions. Compare accepted transcripts, not only the API’s successful responses. Check whether speaker labels remain useful throughout the recording and whether turn boundaries fit the interface. These are proposed evaluation criteria; this publication has not run a production accuracy benchmark on the service.

For a voice-command application, test the endpointing mode instead of borrowing a meeting-demo result. The client should distinguish partial transcripts from completed turns and should not trigger consequential actions merely because plausible words appeared early. Meta’s walkthrough describes asynchronous completion events correlated by turn identifier. That is an integration requirement worth testing before a pricing experiment becomes a user-facing deployment.

Migration has another possible cost: preserving features Meta does not supply. If word-level timing is mandatory, an additional alignment stage might be needed, but this article has no sourced price or performance evidence for such a stage. Do not subtract an invented estimate from the hourly saving and call the remainder profit. Obtain a working design and measured cost, or keep the incumbent for that workload.

The strongest case for switching is an application whose actual requirements match Meta’s turn-level, speaker-aware output. There, included diarization may remove a separately billed component without sacrificing a needed feature. The strongest reason to stay is equally concrete: required timing or confidence information is absent, or representative recordings produce more correction work than the lower rate can justify.

Change the verdict when the full output contract passes, integration behavior remains reliable, and the total cost of accepted transcripts falls. Recheck Deepgram’s promotional rate when pricing the renewal rather than treating today’s gap as permanent. Muse deserves a targeted trial where turn-level attribution is enough; it is not a drop-in replacement for every timestamped transcription pipeline.

Sources