一龍馬/AI 情報站讀懂消息背後的脈絡
星期四
搜尋

Google 新推出的 Gemini 3.8 Live 主打語音直進直出的單一系統,省去語音轉文字再轉語音的中繼延遲

中文摘要

DeepLearning.AI 引述 Artificial Analysis 的語音對語音評測,指 Extended Thinking 版拿下第一,標準版在真人盲測即時對話排第二,且標準版輸入音訊每小時 0.84 美元為評測中最低。兩款模型同時支援圖像與影片輸入,訴求可直接看著螢幕畫面回答問題,但實際延遲與中文語音表現仍待完整評測驗證。

一龍馬判讀

對語音助理與即時客服開發者來說,單一語音模型若兼顧排名與低單價,可能改寫技術選型;採用前仍須確認評測方法、語言支援與總持有成本。

原文節錄

DeepLearning.AI · @DeepLearningAI

Speech-to-speech models skip the relay.…

取得全文 · 不代表內容已獨立查證

查看原文
完整收錄文字與來源

Most voice assistants work like a relay race: speech becomes text, text goes to a model, the answer becomes speech again. Every handoff adds a pause. ⏱️ Speech-to-speech models skip the relay. Google's new Gemini 3.8 Live models listen, reason, and respond in one system. 🎯 The Extended Thinking version ranks first on Artificial Analysis' Speech to Speech Index 🗣️ The standard version ranks second in blind live conversations judged by people 💰 The standard version costs $0.84 per hour of input audio, the lowest in the index Both models also take in image and video input. Picture asking an assistant for help with whatever is on your screen, and it simply answers. 📱 Read the full story in The Batch 👉 https://hubs.la/Q04ztmlL0 #DeepLearningAI #VoiceAgents #AI

收錄日期
2026-10-08
來源
Nitter RSS(公開貼文)
抓取時間
2026/10/08 06:30(台北)