字節跳動 Seed 團隊推出 SeedRealtime:原生視聽全雙工 LLM,單一模型就能看、聽、說

2026年8月10日 05:48
站內 AI 整理稿

ByteDance’s Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM.The model fuses audio, video and text in a single unified architecture.It interacts in real time over continuous multimodal streams, rather than one turn at a time.

Seed positions it as a step toward omni-modal interaction, and claims three breakthroughs: joint audio-visual understanding, proactive interaction, and natural conversational timing.

The architectural target is the cascade: chained ASR, VLM and TTS modules that add latency and lose information between stages.SeedRealtime instead runs perception, understanding, decision-making and expression in parallel inside one end-to-end model.

Turn-taking moves inside the model as well, replacing the external voice-activity detector most real-time stacks still depend on.Is it deployable?It is partly deployable.SeedRealtime is live inside the Doubao app, ByteDance’s consumer assistant.

For this specific model, ByteDance has published no technical report, no parameter count, no open weights, and no Volcano Engine or BytePlus endpoint.As a third-party team, you cannot integrate it as of now.

What is deployable right now is the idea: a validated reference architecture, and a moved goalpost for anyone shipping real-time voice-plus-camera products.Interactive explainer (function(){ window.addEventListener("message",function(e){ if(e.data && e.data.mtpSrHeight){ var f=document.

getElementById("mtp-sr-frame"); if(f) f.style.height=e.data.mtpSrHeight+"px"; } }); })(); What is actually new in the demos Seed published seven scenarios.Four are load-bearing.

Identity binding across modalities: At a noisy group dinner, the model matches names to faces as people are introduced, then keeps each voice tied to its identity — attributing conflicting travel preferences to the right speaker before proposing a plan.

Proactive speech from a held instruction: At the Hebei Museum, a user asks to be reminded when a specific bronze screen stand appears.The camera keeps panning; the model watches and speaks up unprompted when the piece enters frame.

The same behavior shows up on a ResNet paper — the model tracks fast page flips, spots the “3.4 Implementation” section, pauses on its own, and reads out learning rate, momentum and weight decay.

Correction from visual state, not from a question: Watching an espresso workflow, the model interrupts when whole beans go into the portafilter, then reads crema color and volume and suggests shortening extraction by 2 to 3 seconds.

Interference suppression pand off-screen memory: At Beijing Daxing Airport, unrelated chatter about a flight does not trigger a reply.

When the user actually asks, the model answers from departure-board information that has already scrolled off screen, and goes online for the baggage-carousel location.Key Takeaways SeedRealtime is a native audio-visual full-duplex LLM — audio, video and text in one end-to-end architecture.

Turn-taking moves inside the model; no external VAD decides when to speak.ByteDance’s own human eval reports pacing issues halved versus cascaded stacks — no benchmark, no latency numbers.It is live in the Doubao app, but there is no technical report, no weights and no announced API.

Check out the ByteDance Seed launch post and Seed models page.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model appeared first on MarkTechPost.

Related

相關文章

鈦媒體生成式AI

王興興“錯配”梁文鋒?

王興興“錯配”梁文鋒?字母榜2026.08.24 17:44 · 來自河北全文4818字00:00 / 13:21關於世界模型,宇樹和DeepSeek理念分歧明顯。文 | 字母榜宇樹科技的股價還在持續下跌。市值從上市首日的4449億元高點,跌至2400億元,截至8月24日收盤,市值較最高點蒸發了2000億元。然而比股價更值得關注的是,宇樹接下來要怎麼走。8月20日,也就是宇樹上市第二天,王興興出現在北京世界機器人大會論壇。十多分鐘的分享裡,AI成為了高頻詞彙。他談到AI實時生成、實時識別,也談到AI模型投入,更透露了宇樹正在預研的一件事:讓物理AI機器人實現“自進化”。王興興講的每一件事,最後都指向一個關鍵要素:AI大模型。特別是最後一點,王興興說,要實現“物理AI自進化”,要用目前最前沿、最頂尖的AI大模型來驅動。而這恰恰是宇樹目前不太擅長的部分。不過,宇樹找到了DeepSeek。今年8月,兩家公司已經達成合作,圍繞AI大模型與具身智能相關技術展開合作。說到具身智能公司和AI大模型公司的合作,就不得不提當年Figure AI和OpenAI的合作。2024年,OpenAI在投資Figure AI後,雙方簽署了三年合作協議,合作開發人形機器人AI模型。但是僅一年後,FigureAI終止合作,轉向自研。於是問題來了:王興興和梁文鋒,會不會重走Figure AI和OpenAI的老路?01為什麼宇樹需要DeepSeek?對於宇樹來說,過去幾年,模型研發已經有了一些積累,但顯然投入不足,進展緩慢,所以找到一個頂尖的AI大模型公司合作,是一個比較自然的選擇。先來看宇樹目前的模型研發進展。當前,具身智能大模型沒有一個統一的技術路線,但VLA和世界模型,是行業重點探索的兩個方向。行業甚至也在探索兩者融合的技術路線。宇樹也在摸索,採取了“兩條腿走路”的策略,並行研發兩種模型。所謂VLA(視覺

剛剛
量子位生成式AI

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”

阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

剛剛
IT之家生成式AI

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入

作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

剛剛