Google 推出 Gemini 3.8 Flash TTS 與 Flash-Lite TTS,支援提示詞語音設計

2026年9月23日 20:20
站內 AI 整理稿

Google has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, 2 new text-to-speech models in its Gemini Audio family.Google calls them its most expressive audio generation models yet.Flash TTS targets creative direction and character voices.

Flash-Lite TTS targets high-volume, cost-efficient production.Both let developers direct delivery line by line using natural language.Is it deployable?Yes, both models are rolling out now through the Gemini API and Google AI Studio.Access is API-only, with no open weights for self-hosting.

Enterprise API access via Gemini Enterprise is listed as coming soon.What Google Shipped The release splits TTS into 2 tiers with shared direction controls: Gemini 3.8 Flash TTS is built for deep creative direction and character design.

Target uses include gaming, immersive audiobooks, podcasts and interactive media.It offers granular control over acting cues, pacing, dialect shifts and backchanneling.Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale.

Google positions it for dubbing, audio content creation and expressive voice agents.It offers fine-grained control over tone, pacing and expressive nuance.In AI Studio, the playground links use the model identifiers gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.

Voice Design From a Text Prompt Previous Gemini TTS offered 30 original voices.The 3.8 release moves to a much larger voice system.Generative voice design: Flash TTS creates new voices from prompts describing role, accent and voice characteristics.

This works across more than 100 languages and dialects.Google’s demos include a Melbourne DJ, a monotone robot and a Japanese dragon.Voice library: Developers get 2,000+ production-ready voices.Coverage includes regional varieties like Mexican Spanish, Quebec French and Scots English.

Save and scale: Custom voices can be saved and reused, with minimal drift across projects.Voice remixing (coming soon): Users will adjust a library voice’s timbre, pitch, pace and accent through prompts.Directing the Performance Both models accept stage directions written in the script.

Gemini can also steer delivery from natural script cues.Long-form generation: Voice quality, pacing and timbre hold across hours of continuous audio.Native 2-speaker staging: A single script drives a multi-turn conversation with distinct, separated voices.

Vocal bursts: Non-verbal cues like , and add conversational texture.Backchanneling: Active-listening interjections like |mhm| and |yeah| control reaction beats and comedic timing.Voice Replication and Safety Controls Voice replication builds a consistent vocal profile from a 30-second audio sample.

The sample must be your voice or one you have rights to use.Replication requires a verbal consent recording from the voice owner, matched against the reference speaker.Every clip from Gemini Audio models carries a SynthID watermark.This imperceptible mark is embedded directly in the audio output.

Replicated voices also carry C2PA content credentials.Google points to the Gemini 3.8 Audio model card for its broader safety approach.Benchmark Results Google reports these results for the new models: Hume AI Voice Design Benchmark: Flash TTS ranks #1 overall with a score of 71.4, per Hume AI.

Accent modeling: Flash TTS leads with a score of 60.8.Hume AI Overall Quality Index: Flash TTS ranks #1 and Flash-Lite TTS ranks #2.Voice Arena blind preference: Both models take top positions in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.window.

addEventListener('message',function(e){if(e.data&&e.data.type==='mtp-g38tts-height'){var f=document.getElementById('mtp-g38tts-frame');if(f)f.style.height=e.data.height+'px';}}); Key Takeaways Google launched Gemini 3.8 Flash TTS for creative work and Flash-Lite TTS for scale.

Flash TTS designs new voices from prompts across 100+ languages and dialects.Developers get 2,000+ production voices, up from 30 originals.Voice replication needs a 30-second sample plus a matching consent recording.Flash TTS ranks #1 on Hume AI’s Voice Design Benchmark with 71.4.

FAQ What is Gemini 3.8 Flash TTS?It is Google’s text-to-speech model for creative voice design and line-by-line performance direction.It is available through the Gemini API and Google AI Studio.How is Flash-Lite TTS different?

Flash-Lite TTS is optimized for high-volume, cost-efficient workloads like dubbing and voice agents.It ranks #2 on Hume AI’s Overall Quality Index.Can I clone my own voice?Yes, with a 30-second sample and a verbal consent recording.

It is unavailable in AI Studio in several regions, including the UK, EEA and India.Check out the Technical Blog.All credit goes to the researcher of this project.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!

are you on telegram?now you can join us on telegram as well.Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Google Releases Gemini 3.

8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design appeared first on MarkTechPost.

Related

相關文章

全天候科技生成式AI

小米18 Pro價格繼續上探,頂配版破萬元大關

鄭敏芳 發表於 2026年09月23日 15:36 摘要:國產直板旗艦的價格正在進一步向萬元檔靠攏。9月23日晚,小米正式發佈小米18 Pro系列。小...國產直板旗艦的價格正在進一步向萬元檔靠攏。9月23日晚,小米正式發佈小米18 Pro系列。

剛剛
IT之家生成式AI

谷歌推出 Gemini 3.8 Flash / Flash-Lite 文本轉語音模型,每一行臺詞都能精確控制

作者:小泵 責編:小泵 評論: 9 月 23 日消息,谷歌今日宣佈,Gemini 家族新增兩款全新文本轉語音模型,將語音生成從靜態預設轉變為動態創意工作室,能夠幫助創作者、開發者和企業打造更豐富、更具表現力的音頻體驗,同時為 Gemini Notebook 和 Google Vids 等產品改進用戶體驗。

剛剛
IT之家生成式AI

羅福莉官宣小米 MiMo-V3 採用全新架構,核心 HySparse 2 今日發佈

作者:歸瀧 責編:歸瀧 評論: 9 月 23 日消息,小米 MiMo 大模型負責人羅福莉今日發文,宣佈 MiMo-V3 即將採用全新架構。其核心 HySparse 2 今日發佈,帶來更少的預填充、更小的 KV 緩存、更出色的長上下文檢索。在 1M(100 萬)Token 長度下,預填充計算量(FLOPs)降低 5.

剛剛
智東西生成式AI

不造手機,阿里憑什麼做AI手機“Agent底座”?

(公眾號:zhidxcom) 作者 | 李水青 編輯 | 漠影 9月23日報道,剛剛,網信部公告新增榮耀YOYO Claw等3款手機端側生成式AI服務備案。就在昨天,2026雲棲大會現場,榮耀剛宣佈即將發佈的Magic9系列將成為首批搭載阿里Qwen Intelligence的正式機型,榮耀Robot Phone同步支持。 此前蘋果智能國行版備案時,曾傳出引入千問大模型。如今榮耀與千問的合作首先“靴子落地”,AI手機產業分工格局加速變化。

17 分鐘前
IT之家生成式AI

SpaceXAI 智能體 Grok Bot 上線首月,周用戶數突破 40 萬

作者:清源 責編:清源 評論: 9 月 23 日消息,據彭博社 22 日報道,SpaceXAI 的 AI 智能體 Grok Bot 上線約一個月,周用戶數便已突破 40 萬。據倫敦一場活動的演示材料和與會人士透露,截至 9 月 14 日,Grok Bot 用戶數達到 41.

58 分鐘前