Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Microsoft launches streaming transcription and two voice models

Microsoft introduces live transcription and two voice models, with character pricing, gated cloning and important Azure public preview limits.

Listen to this article

Microsoft introduced MAI Transcribe 2 Streaming, MAI Voice 2.1 and MAI Voice 2.1 Flash on October 1. The combined speech launch gives developers a live transcription model and two speech generation options for conversational applications, with introductory transcription pricing of $0.54 per audio hour through the end of 2026.

The release covers both sides of a spoken interaction. Transcribe converts incoming audio into text while a person talks. Voice 2.1 emphasizes expressive narration and speaker consistency, while Flash is designed for applications where fast spoken responses matter.

Transcripts arrive before a sentence ends

Microsoft’s transcription documentation describes incremental results that change as more audio arrives, followed by confirmed segments. Developers can connect through an OpenAI Realtime compatible WebSocket interface or use the Azure Speech SDK to manage streaming and connection behavior.

The Realtime integration guide supports 60 languages with automatic detection. Its provisional text replaces the current unfinished suffix, while finalized text accumulates separately. Applications need to track that distinction so an early interpretation does not become a permanent transcript before the model has finished revising it.

The current Realtime interface also leaves important work to the application. It has no server side speech detection or automatic commit. Clients request a final transcript at a pause or the end of a recording. Sessions last up to one hour and accept mono PCM audio at either 16 or 24 kHz.

Two ways to generate a spoken reply

Both voice models support 23 languages and style controls through SSML. The Azure Speech guide shows developers selecting a prebuilt voice with a model specific suffix. Voice 2.1 targets expressive narration and longer content. Flash prioritizes responsive agents, assistants and interactive support flows.

The Voice 2.1 model card describes long form generation using chunks and carried context to maintain consistency. Microsoft also says a single voice identity can remain consistent across supported languages. These capabilities are aimed at narration and multilingual applications that need a recognizable speaker.

The Flash model card lists streaming output for live dialogue, including interruption handling. It also documents voice matching from reference clips lasting 5 to 60 seconds. Cloning requires Microsoft approval and recorded consent from the voice talent. The release does not make that feature unrestricted.

Pricing uses different billing units

  • MAI Transcribe 2 Streaming costs an introductory $0.54 per audio hour through the end of 2026.
  • MAI Voice 2.1 costs $22 per million characters.
  • MAI Voice 2.1 Flash costs $15 per million characters.

The two voice rates can be compared directly because both charge for characters. Transcription uses incoming audio duration, so its rate measures a different part of the workload. A complete voice agent can incur both speech charges as well as costs for any reasoning model and tools it uses.

Azure access comes with preview limits

Microsoft lists all three models in Foundry and its MAI Playground, with the voice pair also available through OpenRouter. The launch announcement marks LiveKit support as coming soon. For streaming transcription, the current Azure guide lists Sweden Central, Central US and South India as available serving regions, with East US 2 still pending.

The Azure documentation labels the transcription and voice features public previews, provided without a service level agreement and not recommended for production workloads. That qualification matters for teams evaluating customer facing agents. The release supplies components for a voice system, while the application still controls conversation timing, tool use and operational safeguards.

Jordan Reid
Jordan Reid

Jordan Reid is focused on AI tools, agents, developer products, and the way technology changes everyday work. Jordan approaches a launch from the user’s side of the screen. What can it actually help someone finish? The voice is practical, conversational, and skeptical of products that turn a simple job into five new settings. Coverage follows coding assistants, creative software, browser agents, and the workflows around them, with attention to pricing, permissions, setup, and the human work that remains.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile