推出 Olmo-core 3:為大型 MoE 打造的開放、可擴展訓練基礎架構

2026年10月1日 15:01
站內 AI 整理稿

Back to Articles Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs Enterprise Article Published October 1, 2026 Upvote - Kyle Wiggers Ai2Comms Follow allenai 📄 Tech Report | 💻 Code | 🧩 Interactive demo Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system.

Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.

Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs.

MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them.

But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs.

As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input.Olmo-core 3 is built to close that gap.

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B.Total parameter capacity grew from 4.

6B to 47B, while training throughput fell by less than 5%.The same infrastructure has been benchmarked at over one trillion total parameters.Building a training stack around how MoEs actually work Olmo-core has evolved with each generation of Olmo.

Our work on sparse models goes back to OlmoE, which used an MoE architecture with 64 routed experts.Olmo 3, by contrast, used a dense architecture, meaning nearly all of the model was active for every token and its training stack was built around that design.

Olmo-core 3 extends the framework with a training system designed for much larger MoE models.Our earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data.

Olmo-core 3 switches to a system based on distributed data parallelism (DDP).It keeps experts resident on GPUs and routes the relevant data to them, avoiding that repeated weight gathering.NVIDIA’s Megatron-Core is an established option for training large MoEs.

Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation.

In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.

Scaling and optimizing MoE training Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient.

Three techniques determine how the model and its training state are split across hardware: Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool.

Pipeline parallelism splits the model’s layers – the successive stages that transform an input – across groups of GPUs, reducing how much of the model each GPU needs to keep in memory.

A distributed optimizer spreads the optimizer state – the additional data used to calculate and apply updates during training – across GPUs instead of storing a full copy on every GPU.

Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory.Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations.

Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work needed to rearrange it.GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back.

And grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.Finally, Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits.

This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats.

We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts.

With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB.

Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.These techniques and optimizations have to work together.

Speeding up one part of training can create costs elsewhere; faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long.

Olmo-core 3 is built around those trade-offs across the full training process, giving us – and researchers using the open stack – control over how the pieces fit together.

Explore our interactive walkthrough to see how data, expert, and pipeline parallelism work together to scale MoE training—from a single GPU to many.Scaling into the trillion-parameter range We’ve benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs, including a 1.

2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs.Its highest observed throughput was 858 TFLOP/s/GPU—a measure of useful model computation per second on each GPU.

These tests used random routing to measure system performance, rather than the quality of a trained model.We’ve also experimented with DeepEP v2, an alternative way of handling communication between experts across GPUs, reaching a configuration with 2.38 trillion total parameters.

This was a short-capacity test rather than a full training run, so it demonstrates the scale Olmo-core 3 can reach rather than sustained training performance.At these scales, systems performance is only part of the picture.

Our technical report also documents experiments that informed how we train MoEs and measure their performance.For example: A score intended to encourage balanced routing could improve even as the actual workload became less balanced.We call this failure token gerrymandering.

Lowering experts’ learning rates – the size of their training updates – because they process fewer tokens did not improve results in the model family we tested.GPU calculations took different amounts of time when the values being processed changed, even with the same matrix dimensions.

Performance comparisons therefore need matching input values as well as matching shapes.Overlapping communication and computation on separate GPU streams did not always make training faster.

In some tests, it slowed end-to-end execution—a reminder that more overlap does not necessarily mean higher throughput.The report explains these findings alongside the approaches we tested and chose not to adopt.

Built for the next generation of Olmo, open for everyone Olmo-core 3 is the foundation for what we’re building next.Our next-generation Olmo will use an MoE architecture, and we’re aiming for it to be our most capable Olmo yet, trained on our largest dataset and with our longest context window.

The new stack lets us scale beyond our previous MoE work while giving us more flexibility to adapt training as models and hardware evolve.

And it’s fully open—researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system.

That’s part of how we think about open model development—model weights are more useful when the infrastructure and training decisions behind them are open too.

For a deeper look at the systems design, experiments, ablations, and approaches we tested along the way, read our technical report and explore Olmo-core 3 on GitHub.More from this author BenchMIRT: What are LLM benchmarks actually measuring?

27 September 1, 2026 Introducing OlmoEarth embeddings: Custom embedding exports from OlmoEarth Studio for downstream analysis 19 August 12, 2026 Community EditPreview Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images Comment · Sign up or log in to comment Upvote -

Related

相關文章

IT之家模型更新

OpenAI 稱其遭遇有組織蒸餾,將矛頭指向月之暗面

作者:潞源 責編:潞源 評論: 感謝網友 愚公騎馬、咩咩洋 的線索投遞!10 月 1 日消息,當地時間 9 月 30 日,OpenAI 在官網發文,稱其近期遭遇一場有組織的蒸餾活動。OpenAI 在文中表示,該活動最初始於 7 月第一週,符合對抗式蒸餾(adversarial distillation)之特徵,即系統性未經授權地利用某模型輸出或推理過程,幫助訓練、復現或改進另一模型。

剛剛
MarkTechPost AI模型更新

NVIDIA 發布 Kumo Tabular:開放表格基礎模型,單次前向傳播即可預測新資料列

NVIDIA 推出 Kumo Tabular,這是一系列用於分類與回歸任務的新型表格基礎模型(TFM)。若您曾接觸過 TabPFN 或 TabICL,對此架構應不陌生。該模型以已標記的資料列作為上下文,並在單次前向傳播中預測新資料列,無需訓練、無需超參數調整,也無需特徵工程。Kumo Tabular 提供 Small、Medium 與 Large 三種版本,參數量約從 2,800 萬至 2.15 億不等,並透過 NVIDIA 開源的 structured-data-models(SDM)函式庫運行。此模型可直接部署,其權重採用 OpenMDW-1.1 授權,允許商業使用;SDM 程式碼則以 Apache-2.0 授權發布,需搭配 Python 3.11 以上版本與 PyTorch 2.7 以上版本,範例程式以 CUDA GPU 為目標環境。SDM 函式庫的附加價值在於,它是一個專為結構化資料設計的 GPU 原生函式庫。

9 小時前
MarkTechPost AI模型更新

Perplexity 發布 pplx-embed-v2-context-9b-preview:可檢索答案及其佐證的語境嵌入模型

Perplexity Research 與 turbopuffer 合作推出 pplx-embed-v2-context-9b-preview,這是一款專為 RAG 管線設計的語境嵌入模型。每個區塊在嵌入時都會將完整文件納入考量。真正的變革在於訓練信號:模型學習檢索答案以及驗證答案所需的上下文,而非僅擷取單一「黃金段落」。 這款模型是否可部署?可以,作為自架預覽版。權重已以 MIT 授權發布於 Hugging Face。載入時需使用 transformers>=5.4.0 並設定 trust_remote_code=True。目前尚未整合至 Perplexity API。模型卡特別註明,權重與介面可能在不提供向後相容性的情況下變更。 為什麼黃金段落不夠好?RAG 系統會將長文件分割成多個區塊。但某個區塊往往依賴於文件中其他地方定義的實體、標題或概念。語境模型正是為瞭解決此問題而設計。

12 小時前
鈦媒體模型更新

個人Agent,人邁向硅基的第一步?

字母AI2026.09.30 19:17 · 來自北京全文5353字00:00 / 15:05信奉硬件為王的蘋果,天要塌了?文 | 字母AI9月初,Meta的Muse上線,不到兩週登上美國App Store免費榜首,Sensor Tower估算截至9月24日累計下載量超過340萬次。

21 小時前