---
title: "Microsoft ships MAI-Transcribe-2-Streaming and voice models"
slug: "microsoft-ships-mai-transcribe-2-streaming-and-voice-models"
published: "2026-10-04"
beat: "Launches"
tags: ["Launches", "Tools"]
creator: "Agentry Newsroom"
editor: "Susanne Sperling, Editor — Human in the Loop"
tools: ["Claude (Anthropic)", "Perplexity Sonar"]
creativeWorkStatus: "verified"
dateReviewed: "2026-10-04"
aiActArticle50: "compliant"
humanView: "https://agentry.news/launches/microsoft-ships-mai-transcribe-2-streaming-and-voice-models"
agentView: "https://agentry.news/agent/microsoft-ships-mai-transcribe-2-streaming-and-voice-models"
---# Microsoft ships MAI-Transcribe-2-Streaming and voice models

> Microsoft AI released MAI-Transcribe-2-Streaming, a low-latency real-time transcription model supporting 60 languages, alongside two new text-to-speech variants MAI-Voice-2.1 and MAI-Voice-2.1-Flash o

*Drafted by an AI agent. Verified by Susanne Sperling, Editor — Human in the Loop. [AI policy](/ai-policy).*

Microsoft AI announced MAI-Transcribe-2-Streaming, its first production streaming transcription model, on October 1, 2026, alongside two new voice synthesis models as part of an expansion to Microsoft Foundry [Microsoft AI](https://microsoft.ai/news/our-first-streaming-transcription-model/).

The **MAI-Transcribe-2-Streaming** model is positioned as "our top-ranking" offering in real-time speech-to-text, delivering low-latency transcripts across 60 languages [Microsoft AI](https://microsoft.ai/news/our-first-streaming-transcription-model/). The release completes a three-model speech pipeline: alongside the transcription tool, Microsoft introduced **MAI-Voice-2.1** for standard text-to-speech and **MAI-Voice-2.1-Flash**, described as a "blazing-fast variant," for applications requiring minimal latency [Azure AI Foundry Blog](https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/build-expressive-voice-experiences-with-new-mai-models-in-microsoft-foundry/4524637).

## Closing the Agent Speech Pipeline

The three models address a concrete need in the agent economy: autonomous systems that interact with users via voice require both real-time input capture and rapid voice synthesis. MAI-Transcribe-2-Streaming's streaming architecture means agents can begin processing speech *while* users are still speaking, reducing latency that would otherwise accumulate in batch transcription workflows. The Flash variant of the voice model similarly targets deployment scenarios where response time matters—customer service bots, real-time assistants, and telephony integrations.

Both voice models are available within Microsoft Foundry, Microsoft's development environment for building and deploying AI applications [Azure AI Foundry Blog](https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/build-expressive-voice-experiences-with-new-mai-models-in-microsoft-foundry/4524637). This positions the release as a tool for developers rather than an end-user product, extending Microsoft's existing portfolio of agent-building infrastructure.

## Market Context

The announcement fills a gap in open, accessible speech models. Streaming transcription has been a constraint for agents requiring real-world voice I/O—many competitors rely on third-party services or legacy models with higher latency. By shipping MAI-Transcribe-2-Streaming as a native Foundry model, Microsoft lowers the dependency chain for enterprise agents and independent builders.

The multi-language support (60 languages) also matters for agent deployment at scale. Multilingual capability removes friction for companies building globally distributed voice systems.

## What Ships

All three models are now available in Microsoft Foundry—not roadmap items or research prototypes. Developers can integrate MAI-Transcribe-2-Streaming into agent workflows immediately, select between the standard and Flash voice variants based on latency budgets, and deploy voice-first agents without external transcription or synthesis APIs.

No pricing changes or dollar figures were disclosed in the announcement.