Skip to main content
Microsoft AI transcribes 60 languages in 100ms, showcasing advanced speech recognition technology.

Editorial illustration for Microsoft's New AI Transcribes 60 Languages in 100 Milliseconds

Microsoft AI Transcribes 60 Languages in 100ms

• 4 min read

Microsoft AI put out MAI-Transcribe-2-Streaming this week, a real-time transcription model the company says now tops the accuracy rankings on Artificial Analysis. The pitch is speed as much as accuracy: the model handles 60 languages and spits out its first partial result in a little over 100 milliseconds, fast enough, Microsoft argues, for a voice agent to start responding before a caller even finishes a sentence. Pricing during the introductory window runs $0.54 per hour of audio, good through the end of the year.

The release isn't just about transcription. Microsoft also shipped two text-to-speech models alongside it. MAI-Voice-2.1 is built to speak 23 languages while keeping a single voice identity consistent, and with a native-sounding accent in each language rather than an obvious foreign inflection.

A lighter variant, MAI-Voice-2.1-Flash, trades some capability for latency, landing at 150 milliseconds and a lower per-character rate. Both voice models can clone a speaker's voice from just a few seconds of sample audio, which raises its own set of questions about how convincing the output actually is, and how the company is trying to keep that capability from being misused.

Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-time transcription. Microsoft says it ranks first for accuracy on Artificial Analysis. The model transcribes 60 languages and delivers its first partial results in just over 100 milliseconds.

Why this matters

The 100-millisecond figure is the real headline here, not the 60 languages. Voice agents that can start forming a response before someone finishes talking close the gap between "chatbot that waits politely" and something closer to a conversation partner. For developers building on top of this, that latency number matters more than most benchmark scores Microsoft will cite, because it's the difference between a voice product that feels broken and one that feels usable.

The $0.54-per-hour introductory price is worth watching too. Introductory pricing has a way of becoming the baseline people design around, then quietly resetting once the promotion ends, so anyone building a product on MAI-Transcribe-2-Streaming should model costs at a higher rate before committing. Microsoft's "first for accuracy on Artificial Analysis" claim is Microsoft's framing, worth checking against independent benchmarks once third parties get their hands on it.

Pairing this with MAI-Voice-2.1 signals Microsoft wants a full-stack voice pipeline, transcription and speech generation, under one roof. For founders in the voice-agent space, that's a competitor worth tracking closely, not just a feature release.

Common Questions Answered

How fast does Microsoft's MAI-Transcribe-2-Streaming model deliver its first transcription results?

MAI-Transcribe-2-Streaming delivers its first partial results in just over 100 milliseconds, enabling voice agents to start responding before a caller finishes speaking. This speed is crucial for creating natural conversational experiences that feel responsive rather than delayed.

How many languages does MAI-Transcribe-2-Streaming support?

MAI-Transcribe-2-Streaming supports transcription for 60 different languages, making it a versatile solution for global voice agent applications. This multilingual capability allows developers to deploy voice solutions across diverse markets without requiring separate models.

What is the introductory pricing for Microsoft's MAI-Transcribe-2-Streaming model?

The introductory pricing for MAI-Transcribe-2-Streaming is $0.54 per hour of audio during the promotional window. This pricing structure makes the high-performance transcription model accessible for developers building voice agent applications.

Why does the 100-millisecond latency matter more than the 60-language support for voice agents?

The 100-millisecond latency is critical because it enables voice agents to begin formulating responses before callers finish speaking, creating a natural conversation experience rather than an awkward pause. This latency performance is the difference between a voice product that feels broken and one that feels genuinely usable to end users.

How does MAI-Transcribe-2-Streaming rank on accuracy benchmarks?

Microsoft reports that MAI-Transcribe-2-Streaming ranks first for accuracy on Artificial Analysis, the company's accuracy ranking system. This top ranking demonstrates that the model achieves both speed and precision in real-time transcription.

LIVE13:48Business AI Usage Climbs 50% Since July Peak, Ramp Index Shows