微软首个实时语音转文字模型MAI-Transcribe-2-Streaming在38个模型中排名第一,WER仅2.5%且延迟0.13s,支持60语言,定价0.54美元/小时。
Microsoft AI 发布了 MAI-Transcribe-2-Streaming,这是其首款流式语音转文本(STT)模型。该模型于 2026 年 10 月 1 日发布,同期还推出了两款文本转语音模型:MAI-Voice-2.1 和 MAI-Voice-2.1-Flash。Artificial Analysis 将其评为 38 款模型中最终转录和首次部分转录准确度的第一名。该模型面向语音 Agent、实时字幕和听写场景,延迟是决定体验的关键因素。
MAI-Transcribe-2-Streaming 是批处理模型 MAI-Transcribe-2(9 月发布)的实时兄弟版本。它支持 60 种语言的转录,并具备自动、持续的语言检测能力。音频流持续输入,在说话者仍在讲话的同时文本流输出。
模型在接收音频后约 100ms 即可输出首批假设(称为 partials,部分转录)。随着上下文信息的到达,它会修订这些部分转录,然后提交稳定的最终转录文本。因此,Agent 可以在句子进行过程中就开始推理或调用工具。Microsoft 团队表示,内部测试显示词语出现的速度是其最接近竞争对手的 2 倍。
AA-WER Streaming 指数基于约 8 小时的音频数据。混合来源为 AA-AgentTalk(50%)、VoxPopuli(25%)和 Earnings22(25%)。延迟从语音结束(由 SileroVAD 检测)时开始计时。
最终转录:WER 2.5%,语音结束后 0.13s,38 款模型中排名第 1。
首次部分转录:WER 2.5%,语音结束后 0.12s,同样排名第 1。
其他模型表现:Grok Voice Transcribe 2.0 为 2.7%、0.49s。Muse Voice Transcribe 为 3.1%、0.16s。
并非最快:Cartesia Ink-2(外部端点)在 0.07s 内返回最终结果,但 WER 高达 4.0%。
首次部分转录与最终转录准确度相同。这对在说话者结束前就采取行动的 Agent 至关重要。Microsoft 还将该模型置于准确度-延迟 Pareto 前沿上。
MAI-Transcribe-2-Streaming 定价为每小 </think>
MAI-Transcribe-2-Streaming 的定价为每小 </think>
MAI-Transcribe-2-Streaming 定价为每小音频 $0.54。这是截至 2026 年底的introductory 价格。Artificial Analysis 将其标准化为每 1000 分钟 $9.00。批处理版 MAI-Transcribe-2 价格为每小 </think>
音频 $0.10。流式版本 Microsoft 的定价高于 xAI 和 Meta,与 Google 的估算费率大致持平。
Microsoft 提供了两种集成路径。Realtime API 适用于已使用 OpenAI Realtime 兼容 WebSocket 的应用。Azure Speech SDK 处理连接管理、重试和音频流传输。两者都返回中间结果和最终结果。
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列 </think>
该模型也可通过 MAI Playground、Vercel 和 Azure Voice Live 访问。LiveKit 支持列为即将推出。
Microsoft 将其与 MAI-Voice-2.1-Flash 配对以实现完整的语音交互。Flash 以 150ms 端到端延迟生成 45 秒音频,定价为每百万字符 $15。MAI-Voice-2.1 覆盖 23 种语言和 26 个地区,定价为每百万字符 $22。
| 功能 | MAI-Transcribe-2-Streaming | Grok Voice Transcribe 2.0 | Muse Voice Transcribe | Gemini 3.5 Transcribe Live |
|---|---|---|---|---|
| 开发商 | Microsoft AI | xAI | Meta Superintelligence Labs | |
| 发布日期 | 2026 年 10 月 1 日 | 2026 年 9 月 18 日 | 2026 年 9 月 1 日 | 2026 年 8 月 26 日 |
| AA-WER Streaming(最终) | 2.5% | 2.7% | 3.1% | 4.0% |
| 最终结果延迟 | 0.13s | 0.49s | 0.16s | AA 来源未报告 |
| 流式价格 / 小时 | $0.54(intro) | $0.20 | $0.18~$0.54(按 token 计费估算) | — |
| 语言 | 60,持续自动检测 | 数十种,自动检测,支持录制中切换 | 训练 70+,验证 25+ | 85+,自动检测 |
| 流式播客说话人分离 | 未说明 | API 包含(流式未确认) | 是,20+ 说话人 | Live 模式不支持 |
| 接口 | Realtime API(WebSocket)+ Azure Speech SDK | WebSocket | WebSocket + 文件端点 | Live API(WebSocket) |
| 开源权重 | 否 | 否 | 否 | 否 |
来源:Artificial Analysis、9to5Mac(Muse 定价)、DataCamp(Grok 定价)、The Batch(Gemini 定价)。验证于 2026 年 10 月 2 日。
AA-WER Streaming 38 款模型排名第 1:WER 2.5%,延迟 0.13s 出最终结果。
首批部分转录同样达到 WER 2.5%,延迟 0.12s。
60 种语言,持续自动语言检测。
每小时 $0.54 的introductory 价格,高于 xAI 和 Meta。
公开预览,无 SLA,无开源权重。
查看技术详情。所有功劳归功于该项目的研究人员。同时,欢迎在 Twitter 上关注我们,不要忘记加入我们的 15 万+ ML SubReddit 并订阅我们的通讯。另外,你还在用电报吗?现在你也可以加入我们了。
想与我们合作推广你的 GitHub 仓库或 Hugging Face 页面或产品发布或网络研讨会吗?联系我们。