Meta AI 推出 Muse Glimmer:可在單張消費級 GPU 上運行的 300 億參數開放權重代理模型
Meta has released Muse Glimmer, a 30-billion-parameter multimodal model distilled from Muse Spark.It is tuned for always-on local agent workflows, and ships under Apache 2.0.A 30B model normally needs over 55 GB of memory at full precision.
Meta compresses it to roughly 4-bit, then adds block-level speculative decoding so it answers fast enough to sit inside a real agent loop.The result runs on one consumer GPU or a Mac, with no network call.Is it deployable?Yes, the weights are open under Apache 2.0.
The Hugging Face collection carries BF16 weights, GGUF k-quants, ExecuTorch builds, and the DFlash drafter.Self-hosting is the day-one path.Which companies: Solo developers and startups can run it on one 24 GB GPU or an M4/M5 Max Mac.Mid-market teams get on-prem inference without a per-token bill.
Regulated enterprises get an air-gappable agent.Meta advises adding system-level guardrails rather than shipping the model as a bare endpoint.Industries: Healthcare, legal, financial services, defense and public sector, manufacturing, and field service.
These are the settings where data residency, offline operation, or latency rule out a cloud call.Applications: Desktop agents that read screenshots, coding agents, and schema-based function calling.Also document and chart understanding, synthetic data generation, and LLM-as-a-judge evaluation.
(function(){ window.addEventListener("message", function(e){ if(e.data && e.data.mtpGlimmerHeight){ var f=document.getElementById("mtp-glimmer-frame"); if(f) f.style.height = e.data.
mtpGlimmerHeight + "px"; } }); })(); Model and training Muse Glimmer is a dense causal transformer with a dedicated perception encoder.Total parameters are roughly 30B, including the vision tower.Grouped-query attention uses 32 query heads and 2 KV heads.
Attention repeats a [Local, Local, Local, Global] pattern with a 2,048 sliding window.RoPE is applied to local layers only, with theta 500,000.The vision side is a ~1.8B ViT-G/14 perception encoder accepting up to 4,096 visual tokens per image.
Context length is 131,072+, vocabulary is 202,048 tokens, and the knowledge cutoff is January 4, 2026.Input is text and image; output is text.Audio is not supported, and video is processed as individual frames.
Training ran in three phases: Pre-training used logit distillation on Muse Spark’s outputs.Mid-training added longer-context, agent-heavy data with richer reasoning traces.
Post-training combined supervised fine-tuning with on-policy distillation and reinforcement learning across general, reasoning, coding, and agentic domains.Fitting 30B onto consumer hardware At full precision the model needs over 55 GB of memory.
Meta compresses weights to approximately 4-bit precision, which brings the language model under 20 GB.That leaves headroom inside a 24 GB or 32 GB envelope.The KV cache, perception encoder, and drafter share it.Two quantized builds ship.K-Quant-Dynamic targets 32 GB VRAM at 0.2% average degradation.
K-Quant-17GB targets 24 GB VRAM at 1.0%.Degradation is averaged over accuracy metrics across 15 common benchmarks.Generation speed comes from DFlash, a block-diffusion drafter that predicts 16 tokens in one forward pass.The main model verifies the block in parallel.
The drafter uses 5 layers, sliding-window attention at 2,048, and 32 query / 8 KV heads.Meta measured K-Quant-17GB at batch size 1 with greedy decoding.On an RTX 5090, throughput rises from 74.9 to 233.4 tok/s, a 3.1x speedup.Apple M5 Max moves from 26.6 to 50.2 tok/s, and M4 Max from 23.7 to 37.
8 tok/s.Benchmarks Meta compares Muse Glimmer against Gemma4-31B and Qwen3.6-27B in thinking mode.It leads on MCP Atlas at 75.5, against 54.2 and 62.5.It also leads on DeepSearch QA at 74.6, Gaia2 at 43.3, and SWE-Bench Pro at 51.2.Reasoning scores follow: AIME 2026 at 94.7, IFBench at 77.
0, AA-LCR at 80.0.Qwen3.6-27B stays ahead on OSWorld-Verified, 75.6 versus 65.9.It also leads TerminalBench 2.1 at 60.7 and SWE-Bench Verified at 77.2.The pattern is consistent.Muse Glimmer wins on agentic orchestration and reasoning.It trails on computer-use and terminal work.
On safety, Siren AgentDojo attack success rate is 28.4 with utility 94.2.Meta states the model does not meet the Frontier AI definition in its Advanced AI Scaling Framework.It rates chem/bio, cyber, and loss-of-control risk at moderate or lower.Key Takeaways 30B open-weights agentic model, Apache 2.
0, distilled from Muse Spark.4-bit quantization fits it in 24 GB VRAM at 1.0% degradation.DFlash 16-token block speculation gives 3.1x decode speedup on RTX 5090.Beats both comparators on MCP Atlas, DeepSearch QA, and SWE-Bench Pro.Trails Qwen3.6-27B on OSWorld-Verified and TerminalBench 2.1.
Check out the Model weights on HF, Details and Meta Blog.Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.?Connect with us The post Meta AI Releases Muse Glimmer: A 30B Open-Weights Agentic Model That Runs on One Consumer GPU appeared first on MarkTechPost.
Related
相關文章

消息稱蔚來智駕負責人任少卿創立具身智能公司,蔚來將戰略投資
首頁 > IT資訊>人物 消息稱蔚來智駕負責人任少卿創立具身智能公司,蔚來將戰略投資 2026/8/24 18:46:59 來源:IT之家 作者:沁滄(實習) 責編:沁滄 評論: IT之家 8 月 24 日消息,據晚點 Auto 消息,8 月 24 日,蔚來 CEO 李斌在智駕全員會上宣佈,智能駕駛負責人任少卿已創立物理 AI 基礎模型和具身智能獨立公司,蔚來將戰略投資並與之開展合作。知情人士稱,“會議很簡短,沒有提及組織架構變動。”一位參與此次全員會的人士稱,李斌和任少卿在會議上明確表示,任少卿將繼續擔任蔚來智能駕駛業務負責人。報道稱任少卿創立的新公司已經完成註冊,估值達到獨角獸級別。任少卿,1988 年出生,2016 年取得中國科學技術大學與微軟亞洲研究院博士學位。2018 年,任少卿參與創立 momenta,任合夥人兼研發總監;2020 年 8 月加入蔚來汽車,後擔任蔚來自動駕駛研發首席專家、副總裁。IT之家查詢公開信息獲悉,除了在蔚來任職,2025 年 9 月,任少卿加入中國科學技術大學,擔任講席教授、博士生導師和通用人工智能研究所所長,研究方向包括人工智能、世界模型、具身智能、深度學習和 AI for Science 等。相關閱讀:《原 Momenta 研發總監任少卿入職,蔚來自動駕駛研發重心轉回國內》《消息稱蔚來智駕再調整:任少卿直管大模型部門,推進“端到端”量產交付》《中國科大教授任少卿、蔚來汽車 CEO 李斌獲中國人工智能最高獎》 投訴水文 我要糾錯 下載IT之家APP,簽到賺金幣兌豪禮 相關文章關鍵詞:蔚來,智駕,任少卿,具身智能智元聯合長隆打造全球首個具身智能主題樂園,含機器人服務酒店等秦力洪:純電滲透率今年有望超 50%,蔚來不會做增程車蔚來李斌稱中國汽車行業進入“決賽最殘酷階段”,接下來三五年最後的玩家差不多塵埃落定奇瑞捷豹路虎 FREELANDER

偷書要賠15億美元,AI巨頭燒幾百萬本書反而合法了
偷書要賠15億美元,AI巨頭燒幾百萬本書反而合法了藍字計劃2026.08.24 18:01 · 來自廣東全文4120字00:00 / 11:49只要能訓練出更好的大模型,究竟要銷燬多少本實體書,甚至在實體書之外還會消耗多少前AI時代的“資產”,也許就再也沒人關心了。文 | 藍字計劃,作者|Chester一場現代版的“焚書”,正在美國發生。最近幾年,美國二手書市場出現了一個奇怪現象:不少神秘買家,一次性採購成百上千本書,不挑書、不問價,甚至連冷門舊書都照買不誤。隨著電子閱讀的普及,現在還有多少人看實體書?更別說這種成百上千的掃貨始買書。書商們自然也開始好奇:究竟是誰在買走這批書?最近,美國科技媒體404 Media聯繫到一名二手書商。對方剛剛通過書籍交易平臺Biblio,接到了一筆大約1000本舊書的訂單,其中還包括不少稀有、絕版和具有收藏價值的書。為了找出背後的買家,404 Media把一枚AirTag塞進其中一本書裡,然後一路追蹤。最終,這本書被送進了亞馬遜位於拉斯維加斯的一家倉庫。據404 Media調查,這些書到了亞馬遜倉庫後,命運卻只有三步:切掉書脊、批量掃描、然後扔進碎紙機。負責這項工作的VGT3團隊,甚至還給自己設計了一個頗為應景的Logo:一隻露著牙齒、手裡抓著書的霸王龍。 自己也賣書的亞馬遜卻幹起了銷燬實體書的活已經夠驚奇了,但更驚奇的是同樣這樣乾的,不只有亞馬遜。在更早之前,Anthropic就被曝光過一個代號為“巴拿馬計劃”的秘密項目。他們的目標,是把全世界的書都“破壞性掃描”一遍。在大約一年時間裡,Anthropic為此花費數千萬美元,買下數百萬本實體書,然後切掉書脊、掃描內容,再把剩下的紙張送去回收。美國的AI大公司,怎麼就和舊書幹上了?互聯網內容,不夠AI用了AI公司瘋狂買二手書,其實是想買下書裡的內容和數據。過去,互聯網是大模型最方便的數據來源。

阿里視頻大模型Wan3.0正式上線,行業評價“穩定、真實、有質感”
阿里巴巴影片生成大模型Wan3.0正式上線,單次可生成30秒影片,並首次支援doc、xls、ppt、pdf、md等文檔輸入。企業用戶普遍評價其「穩定、真實、有質感」,能穩定保持角色與場景一致性,並已進入短劇、影視、廣告等生產流程。即日起可於阿里雲百鍊、千問等平台體驗,標準版並推出限時7折優惠。

月之暗面第一代萬億參數多模態模型 Kimi K2.5 官宣月底結束服役
作者:歸瀧 責編:歸瀧 評論: 8 月 24 日消息,月之暗面 Kimi 官方微博今日宣佈,其第一代萬億參數多模態模型 —— Kimi K2.5 本月底即將結束服役。據此前報道,今年 1 月,月之暗面宣佈推出並開源了其最新的 Kimi K2.

消息稱知名 AI 研究員 Luke Metz 離開 OpenAI,加入 Meta 超級智能實驗室
作者:遠洋 責編:遠洋 評論: 感謝網友 華南吳彥祖 的線索投遞!8 月 24 日消息,據知情人士向 Axios 證實,知名 AI 研究員 Luke Metz 已加入 Meta 的超級智能實驗室(Superintelligence Labs)。

Anthropic 最強大模型 Fable 5 遇冷,企業用戶轉向更便宜 AI 產品
作者:遠洋 責編:遠洋 評論: 8 月 24 日消息,據英國《金融時報》報道,Anthropic 的美國客戶正在使用更便宜的替代品來替代其最強大的 AI 工具,這在其預計將實現有史以來規模最大的 IPO 之前,對其高支出的商業模式提出了質疑。