Executive Summary
Olmo-core 3 introduces an open and scalable training infrastructure for large Mixture-of-Experts (MoE) models, designed to handle trillion-parameter scales while maintaining computational efficiency. The framework addresses the high cost and communication overhead associated with training large MoEs by redesigning the MoE training system. A key achievement is scaling the expert pool from 8 to 128 while keeping active parameters per token stable, increasing total parameter capacity from 4.6B to 47B in one benchmark without significantly impacting training throughput.
The infrastructure shift involves moving from fully sharded data parallelism (FSDP) to distributed data parallelism (DDP), keeping experts resident on GPUs and routing data directly to them to avoid repeated weight gathering. This change resulted in a 2.7x throughput improvement in a preliminary test involving a 47-billion-parameter MoE. The system incorporates techniques like expert parallelism, pipeline parallelism, and a distributed optimizer for effective distribution. Furthermore, optimizations such as rowwise expert parallelism, GPU-resident routing, and grouped GEMM enhance efficiency by minimizing data movement during the routing process.
The framework also supports lower precision formats like MXFP8, which yielded a 21% higher training throughput in a benchmark while reducing peak active memory. The system’s goal is to provide an open foundation allowing researchers to scale MoE training and adapt infrastructure as models and hardware evolve.
Facts Only
* Olmo-core 3 features a redesigned open Mixture-of-Experts (MoE) training system.
* The system aims to scale MoE training into the trillion-parameter range while preserving computational efficiency.
* In one benchmark, increasing the expert pool from 8 to 128 with four experts per token maintained an active parameter count of about 3.2B per token.
* Total parameter capacity grew from 4.6B to 47B in that test while training throughput fell by less than 5%.
* The infrastructure has been benchmarked at over one trillion total parameters.
* Olmo-core 3 switches from Fully Sharded Data Parallelism (FSDP) to Distributed Data Parallelism (DDP).
* The new system keeps experts resident on GPUs and routes relevant data, avoiding repeated weight gathering.
* A preliminary test on eight NVIDIA B300 GPUs showed a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared to 19,400 with the earlier implementation.
* The system incorporates expert parallelism, pipeline parallelism, and a distributed optimizer for distributing model states across GPUs.
* Rowwise expert parallelism places routed data into expert input buffers to minimize rearrangement work.
* MXFP8 support resulted in 21% higher training throughput compared to BF16 baseline during testing on four NVIDIA B300 GPUs.
Full Take
The narrative centers on shifting the burden of MoE training complexity from expensive, centralized communication strategies (like FSDP) to decentralized, GPU-resident routing and parallelization techniques. The evolution from an MoE architecture focused on dense representation (Olmo 3's design) to sparse scaling necessitates a fundamental change in how parallelism is managed—moving from weight gathering to expert residency. The most significant pattern emerging is the necessary trade-off analysis inherent in scaling: performance gains achieved through faster computation often create new bottlenecks in data movement, as evidenced by the note that overlapping communication did not always improve throughput.
This suggests a resistance to monolithic optimization; achieving scale requires understanding the interplay between memory distribution (Expert parallelism), model structure (Pipeline parallelism), and state management (Distributed optimizer). The subsequent section detailing findings like "failure token gerrymandering" implies that scaling performance must be accompanied by rigorous, potentially counter-intuitive, analysis of routing and learning dynamics, rather than simply maximizing raw throughput. The implicit claim is that open infrastructure facilitates the necessary experimentation to discover these non-obvious constraints on large-scale AI training.
The implication for cognitive sovereignty lies in recognizing that architectural choices are not neutral; they embed assumptions about efficiency and cost that must be explicitly tested. When building scalable systems, the focus shifts from achieving a single metric (throughput) to understanding the system of trade-offs across multiple dimensions—computation, memory, and communication costs. The fact that performance comparisons require matching input values demonstrates that any claim of superior scaling is inherently conditional upon the specific operational context chosen for measurement. What assumptions about cost versus efficiency are being prioritized when developing these open tools?
From the original · Hugging Face Blog
Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.Read the full story at huggingface.co
Sentinel — Human
The text reads like a technical announcement from a research group, blending concrete performance metrics with philosophical arguments about infrastructure design, indicating human authorship within a specialized domain.
