Microsoft ships MAI-Transcribe-2-Streaming and voice models
Microsoft AI announced MAI-Transcribe-2-Streaming, its first production streaming transcription model, on October 1, 2026, alongside two new voice synthesis models as part of an expansion to Microsoft Foundry Microsoft AI.
The MAI-Transcribe-2-Streaming model is positioned as "our top-ranking" offering in real-time speech-to-text, delivering low-latency transcripts across 60 languages Microsoft AI. The release completes a three-model speech pipeline: alongside the transcription tool, Microsoft introduced MAI-Voice-2.1 for standard text-to-speech and MAI-Voice-2.1-Flash, described as a "blazing-fast variant," for applications requiring minimal latency Azure AI Foundry Blog.
Closing the Agent Speech Pipeline
The three models address a concrete need in the agent economy: autonomous systems that interact with users via voice require both real-time input capture and rapid voice synthesis. MAI-Transcribe-2-Streaming's streaming architecture means agents can begin processing speech *while* users are still speaking, reducing latency that would otherwise accumulate in batch transcription workflows. The Flash variant of the voice model similarly targets deployment scenarios where response time matters—customer service bots, real-time assistants, and telephony integrations.
Both voice models are available within Microsoft Foundry, Microsoft's development environment for building and deploying AI applications Azure AI Foundry Blog. This positions the release as a tool for developers rather than an end-user product, extending Microsoft's existing portfolio of agent-building infrastructure.
Market Context
The announcement fills a gap in open, accessible speech models. Streaming transcription has been a constraint for agents requiring real-world voice I/O—many competitors rely on third-party services or legacy models with higher latency. By shipping MAI-Transcribe-2-Streaming as a native Foundry model, Microsoft lowers the dependency chain for enterprise agents and independent builders.
The multi-language support (60 languages) also matters for agent deployment at scale. Multilingual capability removes friction for companies building globally distributed voice systems.
What Ships
All three models are now available in Microsoft Foundry—not roadmap items or research prototypes. Developers can integrate MAI-Transcribe-2-Streaming into agent workflows immediately, select between the standard and Flash voice variants based on latency budgets, and deploy voice-first agents without external transcription or synthesis APIs.
No pricing changes or dollar figures were disclosed in the announcement.