德國黑森林實驗室發佈Flux3 多模態模型:原生音頻生成, 20 秒音視頻同步輸出

2026年7月24日 06:027200 次瀏覽

重點摘要

德國黑森林實驗室發佈多模態基礎模型Flux3,基於Self-Flow架構及專用圖像、視頻、音頻、動作編解碼器,實現物理與數字環境的統一理解與生成。性能碾壓Luma、Runway,首次原生支持音頻生成,可一次性輸出20秒音視頻同步片段,並具備文本/圖像/視頻轉視頻、關鍵幀轉場和多語言對話等功能。

站內 AI 整理稿

德國黑森林實驗室近日正式發表全新的多模態基礎模型 Flux3,該模型採用 Self-Flow 架構,並搭配專為圖像、影片、音訊與動作設計的編解碼器,能夠在物理與數位環境之間實現統一的語意理解與內容生成。

不同於以往僅專注於視覺或文字的模型,Flux3 首次原生支援音訊生成,在不需額外拼接或後製的情況下,可直接輸出最長 20 秒的同步音影片片段,大幅簡化多模態內容的創作流程。

根據官方釋出的資訊,Flux3 在生成品質與效率上已超越業界知名的 Luma 與 Runway 等平台,尤其在音畫同步、動作流暢度與細節還原方面表現突出。

這項突破意味著使用者只需輸入單一提示,即可獲得包含完整聲音與畫面的輸出,而無需分別處理音軌與影像,為影音創作、遊戲開發與虛擬實境等領域開創更直接的應用可能性。

此外,Flux3 還具備多種轉換能力,包括從文字、圖片或現有影片直接生成新影片,以及利用關鍵幀進行精確的轉場控制。

同時,該模型支援多語言對話功能,能夠在生成的內容中融入不同語言的語音與文字互動,進一步拓展跨文化應用場景。

黑森林實驗室強調,Flux3 的設計目標是讓 AI 能夠像人類一樣同時理解並生成真實世界中的多元資訊,為未來通用人工智慧提供更扎實的基礎。

Related

相關文章

MarkTechPost AIAI應用場景

Best Open Speech Recognition (ASR) Models in 2026: WER, Languages, Latency, and License Compared

Open speech recognition stopped being a Whisper monoculture some time in the last twelve months. In March 2026 Cohere released Transcribe, a 2B Apache 2.0 model that took the top of the Hugging Face Open ASR Leaderboard at 5.42% average word error rate. Five weeks later IBM shipped Granite Speech 4.1 2B at 5.33%. Since then ARK-ASR-3B and MOSS-Transcribe-preview-2B have posted lower numbers still. The top of that leaderboard is now separated by less than one WER point. That has a specific consequence for anyone choosing a model: rank is no longer the deciding variable. License, language coverage, streaming support, and cost per audio-hour are. This roundup compares the field on all four. (function(){ window.addEventListener("message", function(e){ if(!e.data || typeof e.data.mtpAsrHeight !

22 小時前