skip to content
The Weighted Average

Developer Tools

MAI-Transcribe-2 Cuts Audio Cost, Not Review Cost

Microsoft launches $0.10-per-hour transcription, 72.2% below MAI-Transcribe-1.5. The real test is correction work and the year-end expiry.

A studio microphone on an adjustable arm against a pink wall
A studio microphone on an adjustable arm against a pink wall. Photograph by Ritupon Baishya

Microsoft launched MAI-Transcribe-2 on September 3 at $0.10 per audio hour, a promotional rate that makes the first decision a trial, not a wholesale migration. Against the $0.36 per hour listed for MAI-Transcribe-1.5 in Microsoft’s model comparison, the new launch price is a derived 72.2% reduction—but it is promised only through year-end.

Reconstructed on September 7, 2026, from records available by September 3, 2026.

The cheap minute has an expiration date

The arithmetic is unusually clean: subtract $0.10 from $0.36, then divide the $0.26 difference by $0.36. The result measures an audio-processing price change, not a reduction in the cost of producing a usable transcript. That distinction should sit beside the purchase order. A transcription system still has to ingest recordings, preserve speaker identity, handle difficult vocabulary, and get consequential errors in front of a reviewer.

Microsoft's new transcription rate is 72.2% below its predecessor

US dollars per audio hour; version 2 promotion ends in 2026

MAI 1.5MAI 2 promo$0$0.1$0.2$0.3$0.4$0.36$0.1
MAI 1.5MAI 2 promo$0$0.1$0.2$0.3$0.4$0.36$0.1
Microsoft AI launch and model comparison · September 3, 2026

Microsoft is not selling price alone. Its September 3 announcement adds speaker diarization, word-level timestamps, keyword biasing, configurable transcription styles, automatic language identification, and code switching. The company reports support for 60 languages and a 5.2% average word-error rate on FLEURS across those languages. These are vendor-reported evaluation results, not a promise that a particular clinic, contact center, or multilingual newsroom will see the same error rate.

The product changes matter because they can remove work from the surrounding pipeline. Speaker labels can make a call searchable by participant; timestamps can connect a disputed sentence to its recording; vocabulary hints can help with a company’s product names. None of those benefits follows from a smaller word-error number alone. A transcript that gets most words right but assigns a commitment to the wrong speaker is not operationally correct.

The evaluation denominators also deserve care. Microsoft’s model page reports 3.4% FLEURS word error for its top 25 languages, while the launch announcement gives 5.2% across 60. Those are different subsets, not contradictory measurements and not figures to swap opportunistically in a business case. The same page lists the predecessor’s top-25 result at 3.7%. That narrower comparison suggests incremental recognition gains accompany a much larger price change; it does not establish uniform improvement across every language.

Speed needs an equally narrow reading. Microsoft says the new model processes audio up to ten times faster than leading competitors, citing Artificial Analysis’s non-streaming evaluation. Non-streaming batch performance is not conversational turn-taking latency. Teams building live assistants should not use a fast file-transcription result as evidence that interruptions, partial hypotheses, or noisy telephone sessions will behave correctly.

This is a useful extension of Microsoft’s move toward its own model stack: a narrow workload can earn adoption without replacing a company’s general-purpose assistant. It also illustrates the other side of Broadcom’s expanding AI supply commitments. Growing infrastructure revenue does not stop a vendor from making one application-layer meter dramatically cheaper. Buyers should distinguish the cost of capacity from the price of a finished service.

Make the reviewer part of the benchmark

The right first adopters are teams with an existing recording corpus and a measurable correction process. Batch captioning, archived-call analysis, and searchable meeting records offer a controlled comparison: run the same files through both versions, keep the originals, and compare the resulting work. Microsoft explicitly offers both a verbatim style that preserves fillers and false starts and a clean style that removes them. Choose the output contract before comparing quality; a cleaner transcript is not necessarily a more faithful one.

Start with acceptance criteria the downstream system can actually use. Does the output preserve names, amounts, negation, speaker turns, and domain terminology? Do timestamp boundaries let a reviewer jump to the disputed phrase? Can a language switch happen without losing the subject of the sentence? Averages can hide precisely the errors that create expensive follow-up work. Test them directly rather than letting the launch leaderboard stand in for local evidence.

The financial threshold is simple even without assuming a wage or a workload volume. The published price difference is $0.26 for each audio hour. Additional review, integration, or operating expense above that amount per hour erases the inference saving. That is a break-even condition, not an estimate that reviewers actually cost so little. Conversely, if the new speaker labels and timestamps reduce review work, the operational saving can exceed the API discount. Measure that effect instead of asserting it.

Access also merits a sober check. OpenRouter lists MAI-Transcribe-2 through Azure at $0.10 per hour, corroborating a purchasing route and the launch rate. A marketplace listing does not settle an organization’s recording-retention rules, regional requirements, or permission to send customer audio to another processor. Procurement and engineering should review those terms before the benchmark corpus leaves its approved boundary.

There are clear reasons to defer a switch. A regulated workflow may depend on a validated output format. A live application may need behavior that a batch benchmark does not measure. A production pipeline may already have correction tools tuned to its existing model. These are not arguments against a pilot; they are reasons to keep the pilot reversible and to score the entire workflow rather than the recognition endpoint in isolation.

Finally, the promotion creates a second decision date. Microsoft’s announcement specifies the end of the year but does not disclose the subsequent rate. Record the expiry in the procurement calendar, avoid extrapolating the introductory price into a durable annual commitment, and ask for the renewal schedule before moving a critical workload. A benchmark win now and a pricing review later are separate approvals.

The verdict is to test promptly and migrate selectively. Evidence that would change it includes worse correction time on representative recordings, missing speaker or timestamp fidelity, an unsuitable processing agreement, or a post-promotion price that removes the advantage. Cheap transcription is valuable. A cheap transcript that nobody can trust is simply another document waiting for a human.

Sources