Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
Google has released Gemini 3.5 Transcribe, a speech-to-text model for real-time voice interfaces and recorded audio.It ships as two endpoints, not one.gemini-3.5-transcribe handles pre-recorded files through the Interactions API.gemini-3.
5-transcribe-live handles bidirectional streaming through the Live API.Google reports average word error rates of 4.0% streaming and 2.6% non-streaming, as measured by Artificial Analysis.Time to final transcription improves 70% over Chirp 3, the previous model.
Automatic detection covers more than 85 languages, including mid-sentence code-switching.The split between the two endpoints is the part worth planning around.They do not share the same feature set, limits, or price.Is it deployable?Yes, but API-only.
There are no open weights and no self-hosted path.This is a managed-service decision, not an infrastructure one.Company level: Any.Solo developers and startups can start on the Gemini API free tier via Google AI Studio.Mid-market teams move to the paid tier for higher rate limits.
The paid tier also guarantees content is not used to improve Google’s products.Regulated enterprises route through the Gemini Enterprise Agent Platform, which adds provisioned throughput, compliance controls, and volume discounts.
Both developer and enterprise tracks are in public preview, so treat production commitments accordingly.Industries: Contact centers and CX platforms, clinical documentation, media captioning and localization, legal and insurance intake, meeting tooling, and voice-driven developer tools.
Applications: Real-time voice agents, live captioning, post-call analytics pipelines, meeting transcription with speaker attribution, dictation, and voice-controlled interfaces.(function(){ window.addEventListener('message', function(e){ if(e && e.data && e.data.mtpG35T && e.data.
h){ var f = document.getElementById('mtp-g35t-frame'); if(f) f.style.height = e.data.h + 'px'; } }); })(); Two API surfaces, two different products The Live API delivers sub-second, continuous transcription.
It emits interiminputtranscription for speculative partials while someone is still talking, then input_transcription when the turn finalizes.Audio goes in as raw 16-bit PCM at 16kHz mono, in 100ms chunks.It supports automatic, hybrid, and manual voice-activity detection.
Ephemeral tokens let mobile and web clients stream without holding an API key.The constraints are real.Live sessions cap at 10 minutes of continuous streaming.Speaker diarization is not supported.Word-level timestamps are not supported.The Interactions API covers what streaming cannot.
It offers speaker diarization, word-level start and end offsets, and custom vocabulary biasing.The vocabulary list takes up to 1,000 terms, with best results below 100.Standard requests accept up to one hour of audio.That drops to 30 minutes once diarization or word timestamps are enabled.
Verbatim and smart are the real design decision Both endpoints expose two modes.verbatim is the default and returns everything, including fillers, repetitions, and false starts.smart removes disfluencies, resolves spoken self-corrections inline, and applies structured formatting.
Google’s own documented example: “Um, so for the meeting, I think we should, uh, invite Alice and, wait no, Bob and Carol.” Verbatim keeps all of it.Smart returns “For the meeting, I think we should invite Bob and Carol.” Smart mode cannot be combined with word timestamps or diarization.
That is the tradeoff to plan around.A readable summary and an auditable transcript are now two different API calls.Performance As measured by Artificial Analysis, Google reports an average word error rate of 4.0% for streaming and 2.6% for non-streaming.
On the multilingual FLEURS benchmark, across a set of top languages and locales, the model reports 5.50% streaming and 5.04% non-streaming.Against Chirp 3, Google’s previous transcription model, time to final transcription improves by 70%.
Language coverage spans over 85 locales with automatic detection and code-switching handled without configuration.Ecosystem The Live API is already wired into LiveKit, Pipecat, Agora, Fishjam, Vercel, and Vision Agents.
On the consumer side, the model powers Rambler on Android, the Gemini app on macOS, and Google Antigravity.Chrome is listed as coming soon.Key Takeaways Two endpoints, not one: streaming trades diarization and word timestamps for sub-second latency.Reported WER is 4.0% streaming and 2.
6% non-streaming, per Artificial Analysis.Smart mode cannot be combined with timestamps or diarization — pick one per call.Blended cost runs about $0.005/min batch and $0.009/min live; no open weights.Hard limits: 10-minute live sessions, 1-hour files, 30 minutes with diarization on.
Check out the Technical details here.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages appeared first on MarkTechPost.
Related
相關文章

小龍們繞過辦公agent
字母榜2026.08.28 12:33 · 來自北京全文4617字00:00 / 12:40DeepSeek除外,做了“半個”agent。文 | 字母榜辦公Agent成了巨頭們的新角鬥場。但在這場熱鬧裡,智譜、月之暗面、MiniMax等小龍們卻相當冷靜,沒有搞出多大動靜。其實小龍們也做了——Kimi在6月發佈了Kimi Work,智譜在3月發佈了AutoClaw,DeepSeek則在8月開源了 DeepSeek Harness。但很顯然,這並不是什麼“辦公Agent”,只是個人桌面而已。

商湯大裝置助力智象未來實現視頻生成業務向國產算力無感遷移
國產算力跑通視頻生成規模化應用 國產算力走向規模商用,如何跨過從“適配驗證”到“真實生產”的關鍵一步,正成為行業面臨的新課題。近期,商湯大裝置與智象未來以視頻生成業務為切入,跑通了從國產算力適配到規模應用的完整鏈路。 對於以圖像、視頻生成模型為核心的AI企業而言,國產化並非簡單地將模型從一種GPU遷移到另一種GPU。

智譜 GLM-5.3-Flash上線,商湯大裝置提供國產算力支持
國產異構助力前沿智能進入普惠時代 8月26日晚,智譜正式上線並開源GLM-5.3-Flash(320B-A18B),這是GLM-5系列的首個原生多模態模型。其總參數320B,能力超過GLM-5.2,在全球權威的Artificial Analysis Intelligence Index(AA綜合智能指數)中取得57分,進入全球前沿模型能力區間,與Anthropic最受歡迎的模型Claude Opus 4.8得分持平。GLM-5.

谷歌推出《Pokémon Sleep》聯名款 Fitbit Air 活動追蹤器
首頁 IT圈 最會買 設置 日夜間 隨系統 淺色 深色 主題色 黑色 投稿 訂閱 RSS訂閱 收藏 軟媒應用 App客戶端 要知App 軟媒魔方 業界 手機 電腦 測評 視頻 AI 蘋果 iPhone 鴻蒙 軟件 智車 數碼 學院 遊戲 直播 5G 微軟 Win10 Win11 專題 搜索 首頁 > 智能時代>智能穿戴 谷歌推出《Pokémon Sleep》聯名款 Fitbit Air 活動追蹤器 2026/8/28 10:21:20 作者:溯波(實習) 責編:溯波 評論: 8 月 28 日消息,Google(谷歌)當地時間 27 日宣佈推出《Pokémon Sleep》聯名款 Fitbit Air 無屏活動追蹤器。

OpenAI的“親兒子”,想要用中國開源模型拿回自主權
Alter2026.08.28 08:17 · 來自浙江全文4602字00:00 / 11:42Harvey走出OpenAI圍城。文 | Alter一個禮拜前,美國法律科技公司Harvey在X上官宣了一則消息:隆重推出基於Kimi K3訓練的首個自有模型Harvey Tenet。消息一齣,迅速在AI圈掀起了一場波瀾。基於開源模型後訓練的事已經屢見不鮮。早在3月份,Cursor的自研編程模型Composer 2,就被扒出是基於Kimi K2.

谷歌推出 Gemini Omni 1.1 Flash 視頻生成 AI 模型,最高 4K 分辨率
作者:沁滄(實習) 責編:沁滄 評論: 8 月 28 日消息,當地時間 8 月 27 日,谷歌宣佈推出 Gemini Omni 1.1 Flash 視頻生成 AI 模型,最高可生成 4K 視頻。附新模型亮點如下:場景擴展:用戶可以直接基於已有視頻,從結尾處無縫銜接並繼續生成後續畫面。