Microsoft launches streaming transcription and two voice models
Microsoft introduces live transcription and two voice models, with character pricing, gated cloning and important Azure public preview limits.

Microsoft introduced MAI Transcribe 2 Streaming, MAI Voice 2.1 and MAI Voice 2.1 Flash on October 1. The combined speech launch gives developers a live transcription model and two speech generation options for conversational applications, with introductory transcription pricing of $0.54 per audio hour through the end of 2026.
The release covers both sides of a spoken interaction. Transcribe converts incoming audio into text while a person talks. Voice 2.1 emphasizes expressive narration and speaker consistency, while Flash is designed for applications where fast spoken responses matter.
Transcripts arrive before a sentence ends
Microsoft’s transcription documentation describes incremental results that change as more audio arrives, followed by confirmed segments. Developers can connect through an OpenAI Realtime compatible WebSocket interface or use the Azure Speech SDK to manage streaming and connection behavior.
The Realtime integration guide supports 60 languages with automatic detection. Its provisional text replaces the current unfinished suffix, while finalized text accumulates separately. Applications need to track that distinction so an early interpretation does not become a permanent transcript before the model has finished revising it.
The current Realtime interface also leaves important work to the application. It has no server side speech detection or automatic commit. Clients request a final transcript at a pause or the end of a recording. Sessions last up to one hour and accept mono PCM audio at either 16 or 24 kHz.
Two ways to generate a spoken reply
Both voice models support 23 languages and style controls through SSML. The Azure Speech guide shows developers selecting a prebuilt voice with a model specific suffix. Voice 2.1 targets expressive narration and longer content. Flash prioritizes responsive agents, assistants and interactive support flows.
The Voice 2.1 model card describes long form generation using chunks and carried context to maintain consistency. Microsoft also says a single voice identity can remain consistent across supported languages. These capabilities are aimed at narration and multilingual applications that need a recognizable speaker.
The Flash model card lists streaming output for live dialogue, including interruption handling. It also documents voice matching from reference clips lasting 5 to 60 seconds. Cloning requires Microsoft approval and recorded consent from the voice talent. The release does not make that feature unrestricted.
Pricing uses different billing units
- MAI Transcribe 2 Streaming costs an introductory $0.54 per audio hour through the end of 2026.
- MAI Voice 2.1 costs $22 per million characters.
- MAI Voice 2.1 Flash costs $15 per million characters.
The two voice rates can be compared directly because both charge for characters. Transcription uses incoming audio duration, so its rate measures a different part of the workload. A complete voice agent can incur both speech charges as well as costs for any reasoning model and tools it uses.
Azure access comes with preview limits
Microsoft lists all three models in Foundry and its MAI Playground, with the voice pair also available through OpenRouter. The launch announcement marks LiveKit support as coming soon. For streaming transcription, the current Azure guide lists Sweden Central, Central US and South India as available serving regions, with East US 2 still pending.
The Azure documentation labels the transcription and voice features public previews, provided without a service level agreement and not recommended for production workloads. That qualification matters for teams evaluating customer facing agents. The release supplies components for a voice system, while the application still controls conversation timing, tool use and operational safeguards.



