Microsoft Launches MAI-Transcribe-2 With 60 Languages and $0.10 Introductory Pricing
Microsoft has launched MAI-Transcribe-2, a speech-to-text model supporting 60 languages, with introductory pricing of $0.10 per hour of audio. The September 3 release expands the company’s in-house AI model range and adds capabilities designed to make recordings easier to turn into usable transcripts.
The price is a limited-time offer through the end of 2026, according to Microsoft’s launch announcement. The company reports an average word error rate of 5.2% on the FLEURS benchmark across 60 languages. That is a vendor-reported evaluation, rather than a promise of the error rate users will see on their own recordings.
The release arrives with a strong value pitch, but the practical story is more nuanced than Microsoft’s claim to have the fastest, most accurate and cheapest transcription model. Independent comparisons and the service’s preview conditions give prospective users important context.
What the independent benchmark shows
The Artificial Analysis non-streaming leaderboard, checked for this article, lists MAI-Transcribe-2 at a 2.0% word error rate and a median speed factor of 410.7. The same table prices it at $1.67 per 1,000 minutes, consistent with approximately $0.10 per hour.
Its results place Microsoft second on the listed accuracy ranking, behind Fun-Realtime-ASR-preview at 1.7%. Nova-3 has the highest listed speed factor, while other offerings have lower listed prices. Those comparisons support a competitive combination of speed, accuracy and cost; they do not establish an unconditional lead in every category.
Benchmark percentages also need to remain attached to their test sets. The 2.0% figure from Artificial Analysis and Microsoft’s 5.2% FLEURS figure describe different evaluations. Treating them as interchangeable would give a misleading impression of what changed or how reliably the model will handle a particular language.
More control over the transcript
Microsoft’s Foundry model listing describes speaker diarization, which separates speakers in a recording, and word-level timestamps for locating individual words in the audio. Keyword biasing is intended to improve recognition of specialized vocabulary, names and abbreviations.
The model also offers a choice between verbatim and clean transcription. Verbatim preserves fillers and false starts; clean removes those disfluencies for a more readable result. Automatic language identification and support for multilingual recordings are also included in the catalog description.
These options matter for different reasons. An editor navigating a long interview may benefit from being able to jump directly to a word. A meeting transcript becomes easier to follow when speech is separated by speaker. A transcript prepared for publication may need different treatment from one used to examine exactly how a conversation unfolded.
For example, removing hesitation can make prose easier to read, but it can also hide something an interviewer wants to preserve. The useful feature is having an explicit choice and checking that the selected output suits the intended purpose.
Public preview brings deployment limits
The Microsoft Learn implementation guide states that the feature is in public preview, without a service-level agreement, and is not recommended for production workloads. That limitation is more consequential for deployment planning than a strong launch benchmark.
The guide requires an Azure subscription and a Microsoft Foundry resource for Speech. It accepts audio files under 300 MB in WAV, MP3 or FLAC format. Developers must enable the enhanced mode and select the intended MAI-Transcribe model; speaker diarization and timestamp granularity are configurable.
For teams assessing the release, this makes a controlled trial a more appropriate first step than assuming it is a direct replacement for an established transcription pipeline.
What the introductory price changes
At the announced rate, 1,000 hours of audio would imply $100 in model transcription charges. That calculation excludes any other infrastructure or workflow costs and does not predict pricing after the offer expires.
The more useful operational measure is cost per transcript that is ready to use. If a recording requires substantial correction, the review effort may outweigh a small difference in processing price. A representative trial should therefore include the accents, background noise, overlapping speakers and specialized words that occur in the organization’s actual recordings.
MAI-Transcribe-2 gives developers another credible option for converting audio into structured, searchable text. Its immediate appeal is the combination of a low introductory rate and practical transcript controls. Whether that combination delivers lasting value will depend on real recording quality, correction effort and the terms that follow the preview.
Feature image: AI-generated editorial graphic featuring Microsoft branding.