MiniMax 發布 MiniMax-Music3:開放權重音樂模型,可從歌詞與結構化描述生成完整五分鐘歌曲
MiniMax released MiniMax-Music3, an open-weights text-to-music model.The model takes two separate inputs: lyrics carrying section tags, and a detailed music description.It returns a complete song of up to five minutes in a single generation, as 32 kHz, 16-bit stereo WAV.
The architecture pairs a Hybrid-LM, an 8B Global LLM with a 0.6B Local LLM, with a continuous synthesis stack built on flow matching and a Flow-VAE.Weights, inference code and three documented serving paths shipped the same day.Is it deployable?
Yes, MiniMax published usable weights, inference code and three documented serving paths on day one, so this is deployable now rather than a research preview.Company level: Solo creators, indie studios and mid-market teams can ship on it directly.
The MiniMax-Music3 Community License permits commercial use, but it requires you to display ‘MiniMax-Music3’ prominently in the product UI, and any organization whose aggregate yearly revenue from those products exceeds US$ 20 million must obtain separate prior written authorization from MiniMax.
Anyone hosting third-party generation must also implement and maintain safeguards against infringing outputs.
Industries: Game development, advertising and brand agencies, short-form video and creator tools, e-learning, podcasting, fitness and wellness apps, retail in-store audio, and music-tech SaaS.
Applications: Background scoring for UGC video, adaptive game and level music, localized ad beds and sonic branding, scratch and demo tracks for songwriters, mood-conditioned playlist generation, and offline batch generation where per-song API cost is the constraint.
The Architecture MiniMax-Music3 combines a hierarchical autoregressive stack with a continuous synthesis path.The training tokenizer uses eight layers of residual vector quantization (RVQ).The first, semantic codebook has 16,384 entries and carries core musical semantics and structure.
The remaining seven acoustic codebooks have 1,024 entries each and encode residual detail.Training optimizes the semantic layer first, then all eight jointly.The Hybrid-LM splits the modeling problem.An 8B Global LLM predicts the first RVQ codebook frame by frame and holds long-range structure; a 0.
6B Local LLM predicts the remaining codebooks within each frame.The model card and license state the Global LLM is initialized from Qwen3-8B; the MiniMax Research post says Qwen3.5-8B, so treat the exact base checkpoint as unsettled.The synthesis stage is the more interesting design choice.
Rather than decoding from discrete RVQ tokens, MiniMax fuses the final hidden states of both LLMs and conditions a 2.4B flow-matching module on them, which maps into a latent space decoded by a 123M Flow-VAE inherited from MiniMax Speech.
At inference the discrete tokenizer decoder is not loaded at all.Two-input control Lyrics carry the words and section tags on their own lines: [Intro], [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge], [Instrumental], [Solo], [Outro].
A separate Structured Caption carries Global Metadata, Vocal Details and Arrangement.MiniMax also ships a music-caption-rewriter agent skill that expands a short description into that three-part format offline.Interactive explainer (function(){ window.addEventListener("message", function(e){ if(!e.
data || typeof e.data.mm3Height !== "number") return; var f = document.getElementById("mm3-embed-frame"); if(f) f.style.height = e.data.mm3Height + "px"; }); })(); Running it Three documented paths.
SGLang-Omni is the reference server; the GitHub page specifies two CUDA GPUs, with GPU 0 running Qwen3 and RVQ autoregressive generation and GPU 1 running flow matching and DAV decoding.
The diffusers modular pipeline fits under 24 GB VRAM at full precision, about 22 GB with automatic CPU offload, and down to 8 GB with leaf-level group offloading.ComfyUI has a native Text to Music template using repacked FP16/INT8 weights from Comfy-Org.
Key Takeaways Open-weights model generating complete five-minute songs at 32 kHz, 16-bit stereo, released August 13, 2026.Hybrid-LM design: 8B Global LLM plus 0.6B Local LLM, feeding 2.4B flow matching and a 123M Flow-VAE.
Synthesis runs on fused continuous hidden states, skipping the discrete tokenizer decoder entirely.Runs on two GPUs via SGLang-Omni, under 24 GB via diffusers, or 8 GB with group offloading.Commercial use allowed with visible attribution; above USD 20M revenue needs written authorization.
Check out the Model on HF and GitHub Repo.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?
Connect with us The post MiniMax Releases MiniMax-Music3: An Open-Weights Music Model Generating Complete Five-Minute Songs From Lyrics and a Structured Caption appeared first on MarkTechPost.
Related
相關文章

王興興“錯配”梁文鋒?
王興興“錯配”梁文鋒?字母榜2026.08.24 17:44 · 來自河北全文4818字00:00 / 13:21關於世界模型,宇樹和DeepSeek理念分歧明顯。文 | 字母榜宇樹科技的股價還在持續下跌。市值從上市首日的4449億元高點,跌至2400億元,截至8月24日收盤,市值較最高點蒸發了2000億元。然而比股價更值得關注的是,宇樹接下來要怎麼走。8月20日,也就是宇樹上市第二天,王興興出現在北京世界機器人大會論壇。十多分鐘的分享裡,AI成為了高頻詞彙。他談到AI實時生成、實時識別,也談到AI模型投入,更透露了宇樹正在預研的一件事:讓物理AI機器人實現“自進化”。王興興講的每一件事,最後都指向一個關鍵要素:AI大模型。特別是最後一點,王興興說,要實現“物理AI自進化”,要用目前最前沿、最頂尖的AI大模型來驅動。而這恰恰是宇樹目前不太擅長的部分。不過,宇樹找到了DeepSeek。今年8月,兩家公司已經達成合作,圍繞AI大模型與具身智能相關技術展開合作。說到具身智能公司和AI大模型公司的合作,就不得不提當年Figure AI和OpenAI的合作。2024年,OpenAI在投資Figure AI後,雙方簽署了三年合作協議,合作開發人形機器人AI模型。但是僅一年後,FigureAI終止合作,轉向自研。於是問題來了:王興興和梁文鋒,會不會重走Figure AI和OpenAI的老路?01為什麼宇樹需要DeepSeek?對於宇樹來說,過去幾年,模型研發已經有了一些積累,但顯然投入不足,進展緩慢,所以找到一個頂尖的AI大模型公司合作,是一個比較自然的選擇。先來看宇樹目前的模型研發進展。當前,具身智能大模型沒有一個統一的技術路線,但VLA和世界模型,是行業重點探索的兩個方向。行業甚至也在探索兩者融合的技術路線。宇樹也在摸索,採取了“兩條腿走路”的策略,並行研發兩種模型。所謂VLA(視覺

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。

阿里雲視頻生成模型 Wan3.0 正式上線,支持單次生成 30 秒視頻、文檔輸入
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,阿里雲消息,今天,視頻生成模型 Wan3.0 正式上線。官方稱,Wan3.0 在生成時長、萬能創作、全能參考以及真實世界還原等維度全面升級,單次可生成 30 秒視頻,並首次支持 doc、xls、ppt、pdf、md 等文檔格式輸入,力求準確還原真實世界。