一龍馬/AI 情報站讀懂消息背後的脈絡
星期三
搜尋

Meta AI 發表 Muse Voice Transcribe,稱其為 Meta Superintelligence Labs 首個即時音訊感知模型

中文摘要

貼文列出的能力包括即時串流語音辨識、可處理 20 位以上說話者的 diarization、endpointing、多語與語碼轉換,並能用語言、關鍵字與上下文 biasing 改善準確度。Meta 也稱它在 Artificial Analysis 的串流語音轉文字與公開 diarization benchmark 排名第一,但貼文未附完整測試條件。

一龍馬判讀

即時 ASR 若同時處理多人分離與語碼轉換,會直接影響會議助理、客服、字幕與語音資料管線。需要留意的是,排名聲稱仍要看語言覆蓋、延遲、成本與繁中場景表現。

原文節錄

AI at Meta · @AIatMeta

Introducing Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs.…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

Introducing Muse Voice Transcribe, the first real-time audio perception model from Meta Superintelligence Labs. Muse Voice Transcribe delivers real-time streaming ASR, diarization with 20+ speakers, and endpointing. It’s multilingual with seamless code-switching and improves accuracy with language, keyword, and context biasing. The model ranks first on @ArtificialAnlys streaming speech-to-text and on public diarization benchmarks.

收錄日期
2026-09-02
來源
Nitter RSS(公開貼文)
抓取時間
2026/09/02 06:12(台北)