Meta AI Introduces MetaRoCE: A Clean-Sheet RDMA Transport Built for AI-Scale Ethernet

2026年8月25日 17:25
站內 AI 整理稿

Training and serving frontier models is now a networking problem as much as a compute problem.Collective operations like all-reduce and all-to-all synchronize thousands of accelerators during training, and the slowest transfer sets the pace for the entire job.

Even small amounts of network friction directly strand significant compute capacity.This week, Meta introduced MetaRoCE.It is described as a clean-sheet RDMA transport protocol purpose-built for AI workloads on commodity Ethernet.The design breaks with standard RoCE on its central assumption.

Standard RoCE expects the network to deliver every frame in order, leveraging PFC and discouraging the packet spraying that provides performance in multiplane and large-scale networks.MetaRoCE instead treats the fabric as lossy and pushes ordering, path selection, and recovery into the NIC.

Meta is releasing the specification, a reference software implementation, and a compliance test suite through the Open Compute Project (OCP) Is it deployable?Not yet, the artifacts possibly ships in October, 2026.

Meta may release the MetaRoCE specification, a DPDK-optimized software reference implementation, and its production compliance framework at the 2026 OCP Global Summit.

Hardware support is early: Meta proved it on AMD Pensando programmable NICs, with additional implementations underway from other vendors.

For now this is a fabric-architecture decision, not a procurement one The problem: the fabric sees packets, the NIC sees intent Meta has scaled clusters to hundreds of thousands of GPUs across multiple data centers and regions.

At that size the network sits in the critical path of every training step.Collective operations like all-reduce and all-to-all synchronize thousands of accelerators, and the slowest transfer sets the pace for the entire job.Standard RoCE is the constraint.

It expects the network to deliver every frame in order, leans on PFC, and discourages the packet spraying that provides performance in multiplane and large-scale networks.

MetaRoCE inverts that: intelligence moves to the endpoint, and the network decomposes into many fine-grained logical paths, each with its own real-time telemetry — per-path RTT, ECN state, and utilization.

This builds directly on Meta’s 2024 RoCE-at-scale work and its broader infrastructure evolution.Six design decisions that matter Out-of-order delivery is the default: Packets are sprayed across many paths and arrive out of order by design.

Every packet carries its own destination, so data is written straight to its final memory location as it lands — no reorder buffer, no head-of-line blocking.Sends carry the match to a posted receive buffer, so a Send lands correctly even when messages ahead of it have not arrived.

Multipathing is native: Each path carries a distinct UDP source port as its ECMP entropy, which the NIC can change at any time to move traffic off a bad route.Because each path keeps its own window and round-trip estimate, the transport can tell congestion from failure and rebalance explicitly.

Loss tolerance replaces losslessness: MetaRoCE treats the fabric as lossy — no PFC, no pause frames.A gap in a path’s 256-bit selective acknowledgment bitvector is evidence of loss rather than reordering, so it triggers retransmission of exactly the missing packet, on the path that lost it.

Congestion control runs from both ends: Sender-driven ECN-based AIMD is combined with receiver-driven fair-share rate hints.

In every acknowledgment the receiver returns the share of inbound bandwidth it allocated to that sender, so senders approach the right speed directly rather than searching for it.Incast resolves in one or two round trips.

Topology independence: MetaRoCE asks the fabric for two things every switch already has: ECN marking and ECMP.

It does not require packet trimming, in-network telemetry, credit-based flow control, or switch-side spraying — which means it also runs over vendor clouds whose configuration you don’t control.

Connection state stops exploding: Traditional RDMA gets more ordering or bandwidth by opening more queue pairs — dozens per node pair — each with a congestion window blind to the rest.

MetaRoCE separates the two: one connection carries many independent ordered streams above and many paths below, under one congestion controller.The numbers Meta implemented MetaRoCE on AMD Pensando programmable NICs.

On a 64-node AMD GPU cluster running RCCL collectives, it was compared directly against RoCEv2 across all-reduce and all-to-all, delivering higher throughput and lower flow completion times.

The resilience result is the core statement: MetaRoCE maintains ~86% throughput at 1% packet loss and continues delivering useful bandwidth even at 10% loss rates, converging gracefully rather than collapsing.

Multiplane validation across 4-plane and 8-plane topologies with up to 4,000 concurrent connections confirmed throughput scales linearly with plane count, and simulated plane failures showed traffic redistributing without application involvement or operator intervention.

Open by design MetaRoCE extends the multi-vendor philosophy that OCP’s Ethernet Scalable Unified Network (ESUN) initiative established for the fabric into the transport layer.

Three artifacts ship: the full spec via OCP, a compliance suite that lets vendors prove their implementations match, and libsoftmetaroce as the authoritative behavioral model for silicon development.

Meta has proven it on AMD Pensando hardware, with additional implementations underway from other vendors.Explainer embed window.addEventListener("message",function(e){ if(e.data&&e.data.mtpFrame==="metaroce"){ var f=document.getElementById("mtp-metaroce-frame"); if(f)f.style.height=e.data.

height+"px"; } }); Key Takeaways MetaRoCE is a clean-sheet RDMA transport that treats Ethernet as lossy — no PFC, no pause frames.Packets spray across paths and write straight to memory; no reorder buffer, no head-of-line blocking.

Holds ~86% throughput at 1% loss on a 64-node AMD GPU cluster running RCCL.Needs only ECN and ECMP from switches, so it runs on fabrics you don’t control.Spec, compliance suite, and libsoftmetaroce land at the OCP Global Summit in October.Check out the TECHNICAL DETAILS here.

Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter.Wait!are you on telegram?now you can join us on telegram as well.

The post Meta AI Introduces MetaRoCE: A Clean-Sheet RDMA Transport Built for AI-Scale Ethernet appeared first on MarkTechPost.

Related

相關文章

下一代生產力,就在豆包工作+飛書裡

AI 正在從「問答助手」走向「生產力核心」,而辦公協作平台正是這波變革最直接的落地場景。當豆包工作與飛書深度整合,新一代生產力工具的樣貌也逐漸清晰:AI 不再只是懸浮在側邊欄的聊天機器人,而是滲透進文件撰寫、會議紀要、任務追蹤與訊息協作的每一個環節,讓知識工作者從繁雜的事務中解放出來,專注於更具創造性的決策。 豆包工作所代表的,是將大型語言模型能力轉化為具體工作場景中的執行力。無論是自動生成結構化文檔、梳理會議重點,還是根據對話脈絡即時補齊背景資料,AI 的介入都在縮短「想法」與「產出」之間的距離。

剛剛

8.86 秒,人形機器人百米競速再次刷新紀錄

作者:沁滄(實習) 責編:沁滄 評論: 8 月 25 日消息,據央視新聞報道,8 月 25 日,第二屆世界人形機器人運動會大型組 100 米複賽,第一組成績再次刷新紀錄。獲悉,繼開幕式上跑出的 9.39 秒,來自北京人形機器人創新中心的天工機器人,跑出了 8.

剛剛

營收47億、淨利暴增477%,重估大普微的稀缺價值

張小夏2026.08.25 12:38 · 來自北京全文3474字00:00 / 10:20AI浪潮重塑全球存儲產業競爭格局,國產廠商不再僅僅是跟隨者。AI算力基建的爆發正在攪動全球存儲產業格局,驅動行業迎來超級週期。一方面,AI領域訓練與推理投資逐漸並重,推理場景相關投入顯著提升,KV Cache(鍵值緩存)等AI場景的需求,加劇了包括企業級SSD在內的存儲需求指數級增長;另一方面,雲廠商通過長協提前鎖量,疊加原廠產能向AI產品傾斜,供給端擴產節奏跟不上需求增速,形成結構性供需緊張,推動存儲行業量價齊升。

6 小時前

Claude Code反超GitHub Copilot登頂第一、90%程序員已用上Agent,最新AI編碼調查報告來了

最新的AI編碼調查報告出爐,結果顯示Claude Code已反超GitHub Copilot,登上開發者最愛用的AI編碼工具第一名。這份報告同時指出,高達九成的程式設計師已經在日常工作中使用AI代理(Agent)來輔助開發,顯示AI輔助程式設計已從早期的嘗試階段,全面進入主流應用的新時代。 這份調查報告的數據來源涵蓋全球眾多開發者,其結果在科技圈引發熱烈討論。長期佔據市場主導地位的GitHub Copilot,憑藉其與GitHub平台的深度整合,曾一度是開發者的首選。

8 小時前

華為鴻蒙智家技術溝通會 8 月 26 日舉行:人車家跨域互聯

作者:汪淼 責編:汪淼 評論: 感謝網友 葛問原 的線索投遞! 8 月 25 日消息,華為今日官宣,華為鴻蒙智家技術溝通會將於 8 月 26 日 10:00 舉行。“技術賦能生態,人車家跨域互聯,鴻蒙智家煥新全場景智慧生活體驗。”華為智慧生活官方 8 月 24 日宣佈,鴻蒙智家迎來秋季煥新升級,新增多款鴻蒙智選新品、6 類快捷指令隨心操控、智慧生活 App 支持燈具 24 小時未關提醒等功能。據此前報道,在今年 6 月的華為 nova 16 系列及全場景新品發佈會上,華為終端 BG CEO 何剛發佈了新一代華為鴻蒙智家。據介紹,華為鴻蒙智家連續四年精裝房市場佔有率第一,而此次新一代鴻蒙智家帶來了“1+3+N”解決方案的全新升級。“1”指家庭大腦:計算中樞 | 連接中樞。“3”指交互方式:包括觸控交互(智能中控屏)、語音交互(小藝管家 6.0)、無感交互(傳感器)。“N”指子系統:包括網絡、影音娛樂、安防方面,具體覆蓋照明、遮陽、用水、能耗、家電、傢俱傢俬、冷暖新風等系統。 投訴水文 我要糾錯 下載APP,簽到賺金幣兌豪禮 相關文章關鍵詞:鴻蒙智家,華為,鴻蒙生態餘承東官宣全新三摺疊即將登場,華為鴻蒙 HarmonyOS 7 | Mate XT 2 及全場景新品發佈會定檔全新華為 FreeBuds 7 悅彰耳機官宣今日開啟預售:半入耳主動降噪、4 色可選華為鴻蒙智家秋季升級:新增 6 類快捷指令操控、智慧生活 App 支持燈具 24 小時未關提醒等華為 FreeBuds 7 悅彰耳機官宣:主動降噪能力全面升級,9 月正式亮相華為出行護航功能適配車型和版本更新,手機端最低要求 HarmonyOS 5.

8 小時前