Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
By Jakub Antkiewicz
•2026-10-01T15:07:29Z
AI2 Releases Olmo-core 3 to Tackle Trillion-Parameter MoE Training
The Allen Institute for AI (AI2) has released Olmo-core 3, a significant update to its open framework for developing large language models. The new version introduces a redesigned system specifically for training large-scale Mixture-of-Experts (MoE) models, an architecture known for its computational efficiency. The release aims to address the substantial communication and memory overheads that can hinder MoE training as models scale, providing a more accessible path for academic and smaller labs to develop models in the trillion-parameter class.
How Olmo-core 3 Optimizes MoE Training
Olmo-core 3 moves away from a Fully Sharded Data Parallelism (FSDP) approach to a Distributed Data Parallelism (DDP) based system, which keeps expert components resident on their respective GPUs and routes data to them. This design choice, combined with several other optimizations, significantly improves performance. In a benchmark on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE model saw a 2.7x throughput increase compared to the institute's previous implementation. The framework's efficiency is built on a combination of techniques:
- Expert Parallelism: Distributes the pool of experts across different GPUs.
- Pipeline Parallelism: Splits the model's layers across groups of GPUs to reduce memory requirements per device.
- Distributed Optimizer: Spreads the optimizer state across GPUs rather than replicating it.
- GPU-resident Routing: Keeps routing metadata on the GPU to avoid CPU-GPU data transfer delays.
- MXFP8 Support: Utilizes a lower-precision number format to reduce computation and data movement, delivering a 21% throughput gain in controlled tests over BF16.
Scaling Open Infrastructure Beyond Model Weights
By open-sourcing its core training stack, AI2 is reinforcing a strategy focused on transparency across the entire model development lifecycle, not just the final model weights. The infrastructure has been benchmarked on configurations as large as a 1.2-trillion-parameter model running across 512 GPUs, demonstrating its capability to handle state-of-the-art scales. This provides the broader research community with tools to not only replicate but also build upon advanced MoE architectures, fostering experimentation in an area of AI development that has largely been dominated by the resources of large, commercial entities.
AI2's release of Olmo-core 3 strategically shifts the open-source battleground from model weights to the underlying training infrastructure, directly challenging the competitive moat of proprietary, highly-optimized internal tooling used by large commercial labs.