字節跳動 Seed 團隊推出 SeedRealtime:原生視聽全雙工 LLM,單一模型就能看、聽、說
ByteDance’s Seed team has introduced SeedRealtime, a native audio-visual full-duplex LLM.The model fuses audio, video and text in a single unified architecture.It interacts in real time over continuous multimodal streams, rather than one turn at a time.
Seed positions it as a step toward omni-modal interaction, and claims three breakthroughs: joint audio-visual understanding, proactive interaction, and natural conversational timing.
The architectural target is the cascade: chained ASR, VLM and TTS modules that add latency and lose information between stages.SeedRealtime instead runs perception, understanding, decision-making and expression in parallel inside one end-to-end model.
Turn-taking moves inside the model as well, replacing the external voice-activity detector most real-time stacks still depend on.Is it deployable?It is partly deployable.SeedRealtime is live inside the Doubao app, ByteDance’s consumer assistant.
For this specific model, ByteDance has published no technical report, no parameter count, no open weights, and no Volcano Engine or BytePlus endpoint.As a third-party team, you cannot integrate it as of now.
What is deployable right now is the idea: a validated reference architecture, and a moved goalpost for anyone shipping real-time voice-plus-camera products.Interactive explainer (function(){ window.addEventListener("message",function(e){ if(e.data && e.data.mtpSrHeight){ var f=document.
getElementById("mtp-sr-frame"); if(f) f.style.height=e.data.mtpSrHeight+"px"; } }); })(); What is actually new in the demos Seed published seven scenarios.Four are load-bearing.
Identity binding across modalities: At a noisy group dinner, the model matches names to faces as people are introduced, then keeps each voice tied to its identity — attributing conflicting travel preferences to the right speaker before proposing a plan.
Proactive speech from a held instruction: At the Hebei Museum, a user asks to be reminded when a specific bronze screen stand appears.The camera keeps panning; the model watches and speaks up unprompted when the piece enters frame.
The same behavior shows up on a ResNet paper — the model tracks fast page flips, spots the “3.4 Implementation” section, pauses on its own, and reads out learning rate, momentum and weight decay.
Correction from visual state, not from a question: Watching an espresso workflow, the model interrupts when whole beans go into the portafilter, then reads crema color and volume and suggests shortening extraction by 2 to 3 seconds.
Interference suppression pand off-screen memory: At Beijing Daxing Airport, unrelated chatter about a flight does not trigger a reply.
When the user actually asks, the model answers from departure-board information that has already scrolled off screen, and goes online for the baggage-carousel location.Key Takeaways SeedRealtime is a native audio-visual full-duplex LLM — audio, video and text in one end-to-end architecture.
Turn-taking moves inside the model; no external VAD decides when to speak.ByteDance’s own human eval reports pacing issues halved versus cascaded stacks — no benchmark, no latency numbers.It is live in the Doubao app, but there is no technical report, no weights and no announced API.
Check out the ByteDance Seed launch post and Seed models page.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post ByteDance Seed Introduces SeedRealtime: a Native Audio-Visual Full-Duplex LLM That Watches, Listens and Speaks in One Model appeared first on MarkTechPost.
Related
相關文章

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。

企業AI最後一公里:三路人馬在此交鋒
鄭敏芳 發表於 2026年08月24日 03:09 摘要:尋找自己的位置 2026年世界機器人大會現場,談到這一輪突然走紅的FDE(前線部署工程師),明略科技CEO吳明輝先把時間往回撥了十多年。“12年前我們就在非常認真地研究。”當華爾街見聞·問及FDE與傳統軟件部署有什麼區別時,吳明輝說,兩者都會進入客戶現場,但今天的FDE需要做得更深:一邊把Agent接進真實業務,一邊把現場形成的能力繼續沉澱回後臺。