Microsoft Introduces New Streaming Transcription Model in MAI AI Model Family
Microsoft Corp. has expanded its MAI artificial intelligence model family with the introduction of its first streaming transcription model. This release is accompanied by two additional models focused on text-to-speech capabilities, aimed at developers seeking to create voice agents that can engage in real-time conversations, much like human interactions.
The most significant of the new offerings is MAI-Transcribe-2-Streaming. This innovative model accepts spoken language through a WebSocket connection and continually updates the transcript while the speaker is talking. Once the individual completes their statements, it confirms the final transcript. Microsoft suggests that this functionality can enhance applications with live captions or allow for the processing of user requests even before they finish speaking.
Available on Microsoft’s Vercel AI Gateway, MAI-Transcribe-2-Streaming is priced at 54 cents per audio hour and supports over 60 languages, automatically detecting the language spoken. The model averages the delivery of its first transcript hypotheses within 320 milliseconds, though Microsoft notes that actual performance may vary based on network conditions and the underlying AI systems generating the responses.
Microsoft is not a newcomer to transcription models; it recently launched the MAI-Transcribe-2, which is priced at just 10 cents per audio hour, significantly less than the streaming variant primarily due to its different processing capabilities. The transcription process for MAI-Transcribe-2 only commences once the speaker has finished talking, distinguishing it from its streaming counterpart, which provides ongoing provisional results during speech.
In addition to transcription, Microsoft has introduced two models that focus on generating speech from written text, essential for conversational AI applications. The models include MAI-Voice-2.1, characterized by its expressive, high-fidelity outputs, and MAI-Voice-2.1-Flash, which emphasizes quicker responses at a reduced cost. According to listings on the Vercel AI Gateway, MAI-Voice-2.1 is available at $22 per million characters while the Flash version costs $15 per million characters, both supporting 23 languages.
This latest model rollout indicates a strategic shift for Microsoft as it begins to move away from relying on transcription models from providers such as OpenAI and Anthropic, despite being significant stakeholders in both companies. In July, it was reported that Mustafa Suleyman, Microsoft AI’s Chief Executive, expressed growing concerns regarding the expenses associated with the models from these firms, prompting him to direct the team to intensify development of the MAI model family with the goal of powering tools like Copilot in applications such as Excel and Outlook. “We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost,” Suleyman stated in a Bloomberg interview.
The introduction of these models equips Microsoft with the necessary resources to develop advanced voice AI agents capable of engaging users in natural, human-like dialogues. Effective voice agents must excel in three areas: comprehend spoken input, reason based on that input, and generate audible replies. The MAI-Transcribe-2-Streaming addresses the comprehension aspect, while the MAI-Voice models provide the spoken output. In the middle lies Microsoft’s robust reasoning model, Mai-Thinking-1, which analyzes the transcriptions to determine the agent’s appropriate responses. By employing three distinct models, developers gain enhanced control over the quality, responsiveness, and costs of their voice agents.

