Google 推出 Gemini 3.8 Live 與 3.8 Live Extended Thinking,專為生產級語音代理打造

2026年9月15日 20:53
站內 AI 整理稿

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced live dialogue models to date.Both are native speech to speech models built for real time voice agents.They extend the Gemini Audio family that Google expanded last month with Gemini 3.5 Transcribe.

The release targets a specific gap: voice agents that can reason and execute tools without breaking conversational flow.Is it deployable?Yes, for API based production use.Both models are live today in the Gemini Live API and Google AI Studio.

They are hosted models, not open weights, so there is no self hosted option.window.addEventListener("message",function(e){if(e.data&&e.data.mtpG38H){var f=document.getElementById("mtp-g38-frame");if(f){f.style.height=e.data.

mtpG38H+"px";}}}); What Google Released The launch covers 2 models with distinct roles.Gemini 3.8 Live is built for scale and cost efficiency.It combines conversational intelligence with fluid dialogue and visual grounding.Gemini 3.8 Live Extended Thinking is built for high complexity tasks.

It adds increased intelligence and multi step reasoning while it speaks.Google positions both as a streamlined alternative to cascaded speech pipelines that chain ASR, an LLM, and TTS.Benchmark Results Gemini 3.

8 Live Extended Thinking takes the #1 overall spot on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6.It leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark.It also scores 97.

7% on Big Bench Audio, a reasoning benchmark for audio models.Gemini 3.8 Live secured second place in the Speech Agent Arena, a human preference evaluation.On ServiceNow’s EVA-Bench, Google reports that the models push the Pareto Frontier for complex workflows.

They balance task accuracy with conversational quality, measured on the Live API on Gemini Enterprise Agent Platform.Capabilities for Developers The Live API exposes 5 core capabilities in the new models: Asynchronous function calling: The model executes API and tool calls in the background.

Audio responses keep streaming to the user while tasks finish.Visual context: The model processes live visual inputs in near real time, so agents can understand what users say and see.

Alphanumeric precision: It accurately parses confirmation codes, claim numbers, and technical data, a common failure point in voice systems.Multilingual support: It automatically detects and transitions between 97 supported languages mid conversation, with accent consistency.

Incremental content updates: It merges real time audio with structured data to return context aware responses.Extended Thinking adds configurable thinking for multi step reasoning in the background.

It reasons and speaks simultaneously, using early verbal cues such as “Let me check that” to acknowledge prompts.It then narrates progress step by step while long running tasks execute.

Google’s demos show the model converting sketches plus voice feedback into working React components and coordinating multi step bookings.Pricing and Ecosystem Both models are priced at $0.005/min for audio input and $0.018/min for audio output.

Google states this estimate is based on $3/1M input tokens and $12/1M output tokens.Developers can also build through Live API integration partners that handle real time media streaming infrastructure.These include Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents.

Google is also partnering with Salesforce, Genspark, and Lumeris, which cite the models’ latency, fluidity, and tool calling.Example apps are available on GitHub.Key Takeaways Gemini 3.8 Live Extended Thinking ranks #1 on Artificial Analysis’ Speech to Speech Quality Index with 82.6.It scores 68.

6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio.Gemini 3.8 Live runs tools and API calls in the background while continuing the conversation.Pricing is $0.005/min for audio input and $0.018/min for audio output via the Live API.

All generated audio carries Google DeepMind’s imperceptible SynthID watermark.Check out the technical details and the developer post.All credit goes to the researcher of this project.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?

Connect with us The post Google Releases Gemini 3.8 Live and 3.8 Live Extended Thinking for Production Grade Voice Agents appeared first on MarkTechPost.

Related

相關文章

量子位生成式AI

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向

無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。

56 分鐘前
IT之家生成式AI

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作

作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

4 小時前
鈦媒體生成式AI

月之暗面遞表之後,Kimi 的成色要被驗算三遍

舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

6 小時前

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"

這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。

8 小時前