Microsoft Launches MAI-Transcribe-2-Streaming for Voice Agents
Microsoft MAI-Transcribe-2-Streaming is now available for developers building live captions and voice agents, returning provisional text while a person is still speaking. Microsoft launched the model on October 1 alongside MAI-Voice-2.1 and a faster Flash version for turning agent responses back into speech.
The release is a material extension of Microsoft’s earlier batch-transcription model, not a rename. Independent evaluator Artificial Analysis ranked the streaming system first for both final and first-partial transcript accuracy, although Microsoft lists it as a public preview without a service-level agreement.
The new audio stack has three components:
- MAI-Transcribe-2-Streaming for live speech recognition
- MAI-Voice-2.1 for higher-fidelity speech generation
- MAI-Voice-2.1-Flash for latency-sensitive voice agents
Microsoft Launches MAI-Transcribe-2-Streaming
The transcription model accepts a continuous audio stream and sends two kinds of results: intermediate hypotheses that change as more audio arrives, and final segments that confirm stable text. That design lets an application display words or begin processing an instruction before the speaker finishes.
Microsoft says the first partial result arrives slightly more than 100 milliseconds after audio is received. The service supports 60 languages and automatic continuous language detection, which can reduce the configuration work needed for multilingual meetings, call centers and assistants.
Developers can connect through an OpenAI Realtime-compatible WebSocket API or the Azure Speech SDK. Both routes return intermediate and final results. Microsoft lists call-center transcription, live captioning, voice-driven interfaces and real-time note taking among the intended uses.
The introductory price is $0.54 per hour of audio through the end of 2026. That price is higher than several streaming rivals tracked by Artificial Analysis, so the commercial question is whether lower correction effort and faster partials offset the difference for a specific workload.
AA-WER Measures Accuracy and Waiting Time
Artificial Analysis measured a 2.5% word error rate for the model’s final transcript, delivered 0.13 seconds after the end of speech. That result ranked first among 38 streaming models in the evaluator’s comparison, ahead of Grok Voice Transcribe 2.0 on both error rate and finalization delay.
The first partial transcript also recorded a 2.5% error rate, arriving after 0.12 seconds. Cartesia Ink-2 returned an earlier partial in the same comparison, but with a higher error rate. Microsoft’s model therefore leads the published accuracy table without being the fastest option in every latency measurement.
That qualification matters because voice systems balance two different costs. Waiting too long makes a conversation feel unresponsive, while acting on an unstable partial transcript can send an agent down the wrong path. Applications will still need interruption handling, confidence checks and a policy for revising work when incoming words change the meaning.
Voice 2.1 Splits Fidelity From Speed
MAI-Voice-2.1 and Voice-2.1-Flash complete the other side of a spoken-agent loop. Both synthesize speech in 23 languages, support granular emotion controls and can match a permitted reference voice from a short recording without separate fine-tuning.
The standard Voice 2.1 model prioritizes fidelity and long-form consistency. Microsoft lists model inference latency at about 550 milliseconds and charges $22 per million characters, positioning it for audiobooks, voice-overs and content where expressiveness matters more than the shortest possible response.
Voice 2.1-Flash targets call centers, interactive voice response systems and assistants. Its listed inference latency is about 45 milliseconds, while the launch post says it can generate 45 seconds of audio with 150 milliseconds of end-to-end latency. Pricing starts at $15 per million characters.
Microsoft says instant voice cloning is gated and limited to authorized, consented voices. Its documentation specifies reference clips between five and 60 seconds and includes licensed prebuilt voices. Those controls are important because low-latency cloning makes impersonation easier as well as enabling consistent branded voices.
Public Preview Limits Production Use
All three models are available through Microsoft Foundry and the MAI Playground, with additional routes including Vercel and Azure Voice Live. The two speech-generation models are also accessible through OpenRouter, while LiveKit support is listed as coming soon.
Availability does not mean the streaming transcription service is production-ready. Microsoft’s documentation labels it a public preview, provides no service-level agreement and explicitly advises against production workloads. Regional capacity, reliability under load and post-preview pricing remain open deployment questions.
Teams evaluating the stack should test accented speech, noisy calls, interruptions and code-switching rather than rely on an aggregate leaderboard. They should also measure the complete loop from captured audio through reasoning and tool use to synthesized response, because a fast transcription model cannot compensate for a slow or unreliable agent behind it.
The launch gives Microsoft a clearer end-to-end proposition for real-time voice software: hear speech early, begin work before the sentence ends and answer through a matching multilingual voice. The preview phase will show whether that combination remains accurate and responsive outside controlled benchmarks.