Microsoft has expanded its range of home AI models, introducing MAI-Transcribe-2-Streaming, its first streaming transcription model. This model transcribes speech in real time, beginning to produce text approximately 100 milliseconds after receiving audio input. It supports 60 languages and automatically detects the spoken language, even if it changes during the conversation. The model continuously analyzes speech, revising its transcription as the sentence progresses. This functionality is suitable for applications that cannot tolerate delays, such as dictation, live subtitling, or voice assistants. The transcription begins before the end of the sentence, with initial hypotheses generated shortly after receiving audio, and these are refined as the sentence continues. The model processes the audio stream as it arrives, rather than waiting for the complete recording, allowing the transcription to follow almost immediately while maintaining flexibility to evolve as the meaning becomes clearer. Alongside MAI-Transcribe-2-Streaming, Microsoft has launched MAI-Voice-2.1 and MAI-Voice-2.1-Flash, two models capable of transforming text into speech. MAI-Voice-2.1 can speak in 23 languages, maintaining the same voice with local accents across languages. The Flash variant prioritizes speed and can generate 45 seconds of audio with an end-to-end latency of around 150 ms. These voice synthesis models are used to build vocal agents that listen, think, and respond aloud, such as in customer service. Microsoft aims to reduce the gaps in voice assistants by integrating MAI-Transcribe-2-Streaming with its Voice models, allowing voice assistants to begin processing requests before the user has finished speaking. This could lead to more natural exchanges and applications such as multilingual assistants, customer service, or voice translation. Microsoft claims that MAI-Transcribe-2-Streaming leads the AA-WER Streaming benchmark from Artificial Analysis, both for provisional and final transcriptions. This ranking is based on the word error rate (WER), which measures the proportion of misrecognized words compared to a reference transcription. However, the benchmark is based on approximately eight hours of audio from three datasets and does not allow comparing the model's performance in each of the 60 supported languages. The claim about precision comes from internal tests, and Microsoft does not name its closest competitor. According to one report, Microsoft states that its transcription appears twice as fast as with its closest competitor, though this figure comes from internal tests. The three models—MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash—are now available through Microsoft Foundry, MAI Playground, Vercel, and Azure Voice Live. MAI-Transcribe-2-Streaming is offered at a launch price of 0.54 dollars per hour of audio until the end of 2026. Regarding voice synthesis, MAI-Voice-2.1 costs 22 dollars per million characters, compared to 15 dollars for the Flash variant. The two Voice models are also available through OpenRouter, with an integration with LiveKit announced soon.