Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model.It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash.Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy.
The model targets voice agents, live captions and dictation, where latency decides the experience.What Microsoft Shipped MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-Transcribe-2, released in September.It transcribes 60 languages with automatic, continuous language detection.
Audio streams in continuously, and text streams back while the speaker is still talking.The model emits its first hypotheses, called partials, just over 100ms after receiving audio.It revises those partials as context arrives, then commits a stable final transcript.
An agent can therefore start reasoning or calling tools mid-sentence.Microsoft team states its internal tests show words appearing 2x faster than its closest competitor.What Artificial Analysis Measured The AA-WER Streaming index uses about 8 hours of audio.
The mix is AA-AgentTalk (50%), VoxPopuli (25%) and Earnings22 (25%).Latency is timed from the end of speech, as detected by SileroVAD.Final transcript: 2.5% WER at 0.13s after end of speech, #1 of 38 models.First partial transcript: 2.5% WER at 0.12s after end of speech, also #1.
Runners-up: Grok Voice Transcribe 2.0 at 2.7% and 0.49s.Muse Voice Transcribe at 3.1% and 0.16s.Not the fastest: Cartesia Ink-2 (external endpoints) returns finals in 0.07s, but at 4.0% WER.The first partial is as accurate as the final transcript.
That matters for agents that act before the speaker finishes.Microsoft also places the model on the accuracy versus latency Pareto frontier.window.addEventListener("message",function(e){if(e.data&&e.data.mtpMaiH){document.getElementById("mtp-mai-frame").style.height=e.data.
mtpMaiH+"px";}}); Pricing MAI-Transcribe-2-Streaming costs $0.54 per hour of audio.This is an introductory price through the end of 2026.Artificial Analysis normalizes it to $9.00 per 1,000 minutes.Batch MAI-Transcribe-2 costs $0.10 per hour.
On streaming, Microsoft charges more than xAI and Meta, and roughly matches Google’s estimated rate.How Developers Integrate It Microsoft documents 2 integration paths.The Realtime API suits apps already using an OpenAI Realtime-compatible WebSocket.
The Azure Speech SDK handles connection management, retries and audio streaming.Both return intermediate and final results.The model is also available in the MAI Playground, through Vercel and Azure Voice Live.LiveKit support is listed as coming soon.Microsoft pairs it with MAI-Voice-2.
1-Flash for full voice loops.Flash generates 45s of audio at 150ms end-to-end latency for $15 per 1M characters.MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per 1M characters.
Comparison: MAI-Transcribe-2-Streaming vs Closest Streaming Competitors FeatureMAI-Transcribe-2-StreamingGrok Voice Transcribe 2.0Muse Voice TranscribeGemini 3.
5 Transcribe LiveDeveloperMicrosoft AIxAIMeta Superintelligence LabsGoogleReleasedOct 1, 2026Sep 18, 2026Sep 1, 2026Aug 26, 2026AA-WER Streaming (final)2.5%2.7%3.1%4.0%Time to final0.13s0.49s0.16sNot reported by AA source citedStreaming price / hour$0.54 (intro)$0.20$0.18~$0.
54 (token-billed estimate)Languages60, continuous auto-detectDozens, auto-detect, mid-recording switch70+ trained, 25 verified85+, auto-detectSpeaker diarization in streamNot statedIncluded in API (streaming not confirmed)Yes, 20+ speakersNot supported in Live modeInterfaceRealtime API (WebSocket) + Azure Speech SDKWebSocketWebSocket + file endpointLive API (WebSocket)Open weightsNoNoNoNo Sources: Artificial Analysis, 9to5Mac (Muse pricing), DataCamp (Grok pricing), The Batch (Gemini pricing).
Verified October 2, 2026.Key Takeaways #1 of 38 on AA-WER Streaming: 2.5% WER at 0.13s to final.First partials score the same 2.5% WER, at 0.12s.60 languages with continuous automatic language detection.$0.54 per hour introductory price, higher than xAI and Meta.
Public preview with no SLA and no open weights.Check out the Technical details.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?
now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis appeared first on MarkTechPost.
Related
相關文章

將語音與音樂作為單一連貫音軌共同生成,AI 音樂製作平臺 Suno 推出 Speech 語音功能
作者:沁滄(實習) 責編:沁滄 評論: 10 月 3 日消息,當地時間 10 月 1 日,AI 音樂製作平臺 Suno 宣佈推出 Speech,Suno 宣稱這是業界首個能夠將語音與音樂作為單一連貫音軌共同生成的音頻模型。據悉,該模型並非像傳統工作流那樣先做 TTS(文本轉語音)再拼接 BGM,而是端到端一體化生成,是首個能單次生成完整融合“語音朗讀 + 原創 BGM”的端到端閉源音頻模型。

OpenAI安全團隊持續地震!負責人離職,三名員工因洩密被開
OpenAI安全團隊再傳人事動盪,負責安全透明度的David Robinson已離職,OpenAI尚未公布繼任者。Robinson團隊主要負責撰寫模型「風險說明書」system card,對外解釋內部安全評估與部署決策。此外,另有三名員工因洩密遭到開除,顯示團隊內部問題持續延燒。

openJiuwen X-Router自演進模型路由技術首發,昇騰親和,Agent越跑越省,實測減少50+%Token消耗
openJiuwen推出X-Router自演進模型路由技術,首發強調與昇騰晶片高度親和。該技術能讓Agent依任務複雜度自動選擇合適模型,避免簡單任務誤用昂貴算力、複雜任務卻交給輕量模型的情況。實測顯示可減少逾50%的Token消耗,使Agent在運行上更省成本、更有效率。

OpenAI 稱其遭遇有組織蒸餾,將矛頭指向月之暗面
作者:潞源 責編:潞源 評論: 感謝網友 愚公騎馬、咩咩洋 的線索投遞!10 月 1 日消息,當地時間 9 月 30 日,OpenAI 在官網發文,稱其近期遭遇一場有組織的蒸餾活動。OpenAI 在文中表示,該活動最初始於 7 月第一週,符合對抗式蒸餾(adversarial distillation)之特徵,即系統性未經授權地利用某模型輸出或推理過程,幫助訓練、復現或改進另一模型。
推出 Olmo-core 3:為大型 MoE 打造的開放、可擴展訓練基礎架構
今日我們正式釋出 Olmo-core 3,這是我們大型語言模型開發框架的重大升級,核心亮在於重新設計的開放式混合專家(MoE)訓練系統。Olmo-core 3 旨在將 MoE 訓練規模擴展至兆級參數,同時維持運算效率。該系統是下一代 Olmo 的核心基礎之一,亦體現我們持續開放每款新模型背後工具與訓練基礎架構的承諾。

流式轉錄 AI 新標杆:微軟 MAI-Transcribe-2-Streaming 登場,2.5% 詞錯誤率、0.13 秒延遲奪冠
作者:故淵 責編:故淵 評論: 感謝網友 xxy171070 的線索投遞!10 月 2 日消息,微軟昨日(10 月 1 日)發佈公告,宣佈推出其首個實時流式語音轉寫模型 MAI‑Transcribe‑2‑Streaming,可在講話進行時持續輸出文字,覆蓋 60 種語言,並支持自動語言檢測。