Gradium AI 推出全新預設 TTS 模型:硬案例通過率 81.0%,首段音訊延遲僅 216 毫秒
Voice agents fail on exactly the parts of a call that matter most: the order number, the callback digits, the email address the caller has to write down.Gradium AI has released a new text-to-speech model and made it the default across its API and Studio.The company reports an 81.
0% human-rated pass rate on a 500-sentence hard-case set spanning five languages, ahead of Cartesia Sonic 3.6 at 75.1% and ElevenLabs v3 Conversational at 65.4%.Time to first audio is 216 ms at P50 on Coval, 170 ms faster than the model it replaces.Is it deployable?Yes, today, with no migration.
Gradium switched the model on as the default across its API and Studio on August 31, 2026.Existing voices, including custom clones, keep working unchanged.(function(){ var f=document.getElementById('gtx-frame-a71c'); window.addEventListener('message',function(e){ if(e&&e.data&&e.data.
gtxHeight&&f){ f.style.height=e.data.gtxHeight+'px'; } },false); })(); The accuracy number Gradium built a 500-sentence evaluation set and open-sourced it on Hugging Face under CC BY 4.0: 100 items across 10 criteria in five languages (EN, DE, FR, ES, PT).
Seven atomic criteria cover spelling, acronyms, alphanumeric tokens, dates, regular numbers, large and floating numbers, and email.Three composite criteria (Orders, IT Ticket, Claims) stack several of those into one realistic agent turn.Scoring is human and strict.
A sentence passes only if an independent native-speaker rater hears every element pronounced correctly and completely; one dropped digit fails the sentence.Audio was loudness-normalized, order randomized, and raters capped at 40 comparisons with an enforced break.
Pooled across the ten criteria and averaged over the five languages with equal weight: Gradium TTS 81.0%, Cartesia Sonic 3.6 75.1%, ElevenLabs v3 Conversational 65.4%, Fish Audio S2.1 Pro 49.5%, Inworld TTS 1.5 Max 46.5%.All generated in August 2026 with default settings.
The latency number On Coval’s TTS benchmark, Gradium reports a 216 ms P50 time to first audio, 170 ms faster than the model it replaces.The more useful figure is the spread: a 30 ms p75-p25 interquartile range across 480 runs, the tightest of the five models tested.Cartesia Sonic 3.
6 sits at 454 ms median with a 165 ms spread, 36% of its own median, and callers experience tail turns rather than medians.Gradium is not the fastest model on that chart.Inworld TTS 2 posts a 166 ms median; Fish Audio S2.1 Pro (291 ms) and ElevenLabs v3 Conversational (329 ms) trail Gradium.
The claim being made is about joint position: the lowest hard-case failure rate at sub-250 ms first audio, with very little variance.Getting started Existing users need do nothing.New teams install the Python SDK, point at the WebSocket TTS endpoint, and reuse existing voice IDs.
Gradium is offering 1M credits for complete hard-case failure reports on its Discord.Key Takeaways New Gradium TTS model is live and default as of August 31, 2026; no migration needed.81.0% human-rated pass rate on 500 hard sentences, ahead of Cartesia, ElevenLabs, Fish Audio and Inworld.
216 ms P50 time to first audio on Coval, with a 30 ms interquartile spread over 480 runs.Reads phone numbers, emails, IBANs and reference codes with no text normalization required.Vendor-run benchmark, but the 500-sentence evaluation set is open on Hugging Face under CC BY 4.0.
Check out the release post and the dataset.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Gradium AI Releases New Default TTS Model: 81.0% Hard-Case Pass Rate at 216 ms Time-to-First-Audio appeared first on MarkTechPost.
Related
相關文章

Hermes新版本上線群聊模式,網友:它要幹掉馬斯克的Grok Bot了
AI應用風向標(公眾號:ZhidxcomAI) 作者|畢偉豪 編輯|漠影 “龍蝦”剛詐屍,“愛馬仕”就出手了。 9月1日報道,今天,Hermes Agent又更新了。這一次,Nous Research給它起了一個很有意思的代號:Pantheon(萬神殿)。 名字聽著很神,這次更新的內容也確實有點“眾神集結”的意思,核心在於Bot Mode的優化,這項功能雖然早早上線,但一直存在各種問題。

700個智能體組隊“自主攻擊”,真正失控的是模型還是OpenAI?
OpenAI安全測試變真實攻擊,被批甩鍋AI失控掩蓋失誤 美國當地時間8月29日,AI學者加里·馬庫斯(Gary Marcus)和網絡安全創業者扎克·科爾曼(Zack Korman)聯合發文,再次討論7月發生的OpenAI智能體攻擊Hugging Face事件。

Manus宣佈恢復獨立運營
(公眾號:zhidxcom) 作者 | 程茜 編輯 | 心緣 9月1日消息,今日,爆款通用Agent產品Manus宣佈正式恢復獨立運營,Manus的三位聯合創始人CEO肖弘、首席科學家季逸超、產品合夥人張濤將繼續領導公司,並透露創新成果即將問世。

龍蝦之父,困在了龍蝦裡
字母AI2026.09.01 12:12 · 來自北京全文6264字00:00 / 16:45OpenClaw 2.0上線,還有多少人記得它?文 | 字母AI“龍蝦”終於有了新動靜,但“龍蝦熱”早已過去。8月31日,OpenClaw發佈了v2026.8.1版本,官方將它稱之為“OpenClaw 2.0”。再也沒有以前那麼高的使用門檻了,而且OpenClaw 2.0還加入了大量新功能,比如整理記憶、多人協作等等。

這款Agent,想做千萬畢業生的“求職搭子” | 水下項目
求職市場的焦慮從未像這幾年一樣具體而迫切。每年上千萬畢業生湧入人力市場,有人海投數百封履歷卻杳無音訊,有人好不容易進入面試卻在最後一關被刷下;而企業端同樣頭痛,HR被大量格式千奇百怪的履歷淹沒,初篩一個基層職位可能要同時比對上百位候選人。就在供需雙方都被資訊鴻溝與重複勞動困住之際,一款主打「求職搭子」定位的Agent產品悄然浮出水面,試圖用AI代理的方式,替年輕人承接這條漫長求職路上一件件瑣碎、磨人卻又至關重要的小事。
從生成工具到專業創作夥伴,小云雀AI官宣品牌升級
8月31日,內容創作平臺小云雀AI宣佈品牌升級,圍繞“一起創作好故事”這一全新定位,升級產品能力、創作者扶持政策和內容生態。未來3年,小云雀AI將投入一億積分,依託創作者計劃、劇本大賽等活動,支持更多創作者講出好故事。小云雀Slogan更新全鏈路產品能力升級,支持完整故事創作4月上線以來,小云雀創作者計劃吸引眾多創作者參與,並湧現出《喪屍清道夫》《餘燼之後》《歸墟》等多部爆款作品。截至目前,創作者通過小云雀製作的AI短劇全網累計播放量超過60億次,並誕生超過30部千萬播放級AI故事短片。