Meta AI 推出 Muse Code(Beta 版):由全新 Muse Spark 1.2 模型驅動的終端編碼代理
重點摘要
Meta AI 已發布 Muse Code(Beta 版),這是一款基於全新 Muse Spark 1.2 模型的終端編碼代理。Meta 將此組合定位為邁向技術前沿的下一步,未來還將推出更大規模的模型。Muse Code 專注於大型程式碼庫中的複雜軟體工程:它能夠規劃變更、編寫程式碼並驗證結果。一組非同步背景代理會在整個工作階段中保持活躍,而非每次任務重新生成。本機僅允許附加的事件日誌會記錄每次模型呼叫、工具執行、核准與編輯,Meta 稱其可精確重播且重啟安全。Muse Spark 1.2 是與該 harness 共同訓練的。Meta 還發布了一個核心優化案例研究,在長達 24 小時內執行超過 1,000 次工具呼叫。是否可部署?是的。Muse Code 以 Beta 版形式提供給 macOS 和 Linux 使用者,透過 vi 下載。
Meta AI has released Muse Code (in beta), a terminal coding agent in beta, powered by its new Muse Spark 1.2 model. Meta positions the pair as its next step toward the frontier, with larger models on the way. Muse Code targets complex software engineering across large repositories: it plans changes, writes code, and validates the results. A set of async background agents stays alive for the whole session instead of spawning per task. A local append-only event log records every model call, tool run, approval, and edit, which Meta calls replay-exact and restart-safe. Muse Spark 1.2 was co-trained with the harness itself. Meta also published a kernel-optimization case study running 1,000+ tool calls over as long as 24 hours. Is it deployable Yes. Muse Code ships in beta for macOS and Linux via curl -fsSL https://dev.meta.ai/install.sh | bash. Muse Spark 1.2 is available in Muse Code and the Meta Model API, with expanded global access. The launch post does not mention downloadable weights, so treat this as a hosted dependency. Company level: The API path fits any size. The Muse Code path fits teams already running agents in sandboxes with review gates. Industries: Software and SaaS, developer tooling, fintech engineering, GPU and inference infrastructure, semiconductors and HPC. Applications: Repository-scale refactors and migrations, long-running bug triage, test generation, and GPU kernel optimization. Async background agents Muse Code runs a simple agent loop plus a set of async background agents. These specialized agents remain active throughout each session. They are not spawned for individual tasks, which Meta says avoids redundant information gathering. They carry out next steps and choose when to report back to the main agent. Meta states this persistence reduces latency and steering on difficult, multi-step tasks. Runtime design Muse Code uses a local event log. Every model call, tool run, approval, and edit is appended to it. Meta calls this single source of truth replay-exact and restart-safe. After a crash, the agent resumes precisely where it stopped, letting long-running tasks survive failures. Bundled skills Three default skills ship with the agent. /plan turns a task into an approval-gated plan. /grill stress-tests that plan until it holds up. /goal works toward successful completion of the specified objective. What changed in Muse Spark 1.2 Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1. Meta reports gains in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. The research team significantly scaled up training compute on coding tasks and expanded environment diversity. The model keeps its strength in other areas, including general agents. Three important training details: Co-training with the harness: Muse Spark 1.2 was co-trained with Muse Code. Training included rejection-sampled harness trajectories and recipe optimizations for goals, compaction, and subagents. The Muse Code toolset was integrated to maximize harness compatibility. Long-horizon training: Training covered whole-repository generation, large end-to-end projects, and auto-research. The model uses planning, goal conditioning, and context compaction to sustain progress. Self-improvement: Muse Spark 1.1 generated challenging coding environments and instruction-following templates, then graded candidate solutions against those requirements. That produced a scalable training dataset for 1.2. https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2 Evaluation Meta’s methodology report is unusually specific. Terminal-Bench 2.1 uses all 89 tasks, pass@1 over five attempts. DeepSWE v1.1 covers 113 tasks across 91 repositories and five languages. Meta Internal Coding Bench holds 440 tasks derived from real internal pull requests. Runs execute in isolated Daytona cloud sandboxes. Comparisons include Grok 4.5, Claude Opus 5, GPT-5.6 Terra, Gemini 3.6 Flash, and Kimi K3, each with its own agent product. Meta notes its harness may not be tuned for third-party models. For reference, Meta’s model page lists Muse Spark 1.1 at 80.0 on Terminal-Bench 2.1. (function(){ var f=document.getElementById("mtp-muse-frame"); window.addEventListener("message",function(e){ var d=e.data||{}; var h=d.mtpHeight||d.height; if(!h||typeof h!=="number")return; if(d.type&&d.type!=="mtp-resize"&&!d.mtpHeight)return; if(h>200&&h<6000){f.style.height=h+"px";} },false); })(); Case study: kernel optimization Meta tested iterative GPU kernel optimization over 1,000+ tool calls, running up to 24 hours. The model writes, compiles, profiles, and progressively improves kernels against a provided baseline. Benchmarks covered KDA and MLA kernels on NVIDIA Hopper GPUs. For KDA, the baseline is the FLA Triton implementation, with third-party kernel libraries prohibited. Muse Spark 1.2 paired a chunk-parallel preparation kernel with a sequential inter-chunk scan. For MLA, the reference is PyTorch at batch size 1, 64 heads, sequence length 8192, and latent dimension 512. The model built a two-kernel Triton pipeline that reuses the shared KV latent as both K and V. Key Takeaways Muse Code is a beta terminal coding agent for macOS and Linux, powered by Muse Spark 1.2. Persistent async background agents replace per-task spawning to cut redundant information gathering. An append-only local event log makes the runtime replay-exact and restart-safe after crashes. Muse Spark 1.2 was co-trained with the harness and trained on long-horizon, repository-scale work. Kernel case study ran 1,000+ tool calls over 24 hours on NVIDIA Hopper KDA and MLA kernels. Check out the Technical details, Model (Muse Spark 1.2) and Evaluation Methodology. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us The post Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model appeared first on MarkTechPost.
Related
相關文章

聽說一些公司開始做員工skills了
賬號設置我的關注我的收藏申請的項目退出登錄登錄搜索36氪Auto數字時氪未來消費智能湧現未來城市啟動Power on36氪出海36氪研究院潮生TIDE36氪企服點評36氪財經職場bonus36碳後浪研究所暗湧Waves硬氪氪睿研究院媒體品牌企業號企服點評36Kr研究院36Kr創新諮詢企業服務核心服務城市之窗政府服務創投發佈LP源計劃VClubVClub投資機。

AI替你自動下單買東西,接受嗎?
AI智能體未來可能取代消費者自動下單購物,支付角色也將從交易工具轉變為價值調度中樞。然而,這項技術面臨信任與隱私挑戰,使用者需確保AI能準確理解偏好與安全需求。目前該技術仍在早期階段,普及與否取決於社會能否建立完善的信任機制與法規保障。

Liquid AI 發佈 LFM2.5-2.6B 端側小模型,支持智能手機本地運行
首頁 > 智能時代>人工智能 Liquid AI 發佈 LFM2.5-2.6B 端側小模型,支持智能手機本地運行 2026/8/5 13:54:26 來源:IT之家 作者:溯波(實習) 責編:溯波 評論: IT之家 8 月 5 日消息,人工智能初創企業 Liquid AI 當地時間昨日發佈 LFM2.5-2.6B 開源智能體模型。僅 26 億的參數使得 LFM2.5-2.

上訴法院推翻原裁決,Perplexity 智能體重獲亞馬遜電商訪問法律許可
美國聯邦第九巡迴上訴法院於當地時間8月4日作出裁決,推翻了一審法院針對人工智慧新創公司Perplexity AI所發出的臨時禁令,這意味著該公司的AI智能體將可重新合法訪問亞馬遜電商平台,並繼續為用戶提供AI輔助購物服務。這項二審判決被外界視為AI產業在與大型電商平台法律攻防戰中的一次重要勝利。 這起訴訟的源頭,來自於Perplexity AI推出的Comet瀏覽器及其內建的AI代理功能。該功能允許用戶透過自然語言指令,讓AI智能體在亞馬遜等電商網站上進行商品搜尋、比價、下單等操作。

Cloudflare Wallet 推出:專為 AI 智能體設計的可編程錢包
首頁 > 智能時代>人工智能 Cloudflare Wallet 推出:專為 AI 智能體設計的可編程錢包 2026/8/5 9:26:09 來源:IT之家 作者:故淵 責編:故淵 評論: 感謝IT之家網友 我家小妞很拽 的線索投遞!

4種AI協作模式,守住屬於管理者的核心價值
本文探討管理者如何在AI協作中保留自主能動性,提出四種協作模式,包含領導者主導、領導者塑造、AI輔助與全權處理。核心在於管理者需守住判斷力、表達話語權與在場影響力,避免過度依賴AI導致思維趨同與決策能力退化。