Ai2 releases Olmo core 3 for large MoE training
Ai2 releases an open training stack with larger expert pools, detailed performance tests and migration changes for teams building sparse models.

Ai2 has released Olmo core 3, an open training framework designed to make large mixture of experts models more practical to train. Its October 1 announcement describes a redesigned system that keeps experts on GPUs and moves the relevant token data between them, reducing repeated movement of model weights.
The clearest result concerns stored capacity. In one benchmark, the expert pool grew from eight to 128 while four experts remained active per token. Total parameters increased from 4.6 billion to 47 billion, with roughly 3.2 billion active per token. Throughput fell by less than 5 percent.
Why sparse models still need careful infrastructure
The technical report separates the arithmetic used for each token from the cost of storing and updating the complete model. A sparse model can activate only a few experts while still carrying substantial memory, communication and optimizer overhead across its full expert pool.
The redesign combines distributed data parallelism with expert parallelism, pipeline parallelism and a distributed optimizer. Experts stay resident in their assigned GPU memory. Token rows travel to the relevant experts, while layers and optimizer state are divided across the cluster. Keeping routing metadata on the GPU removes recurring waits on the CPU.
For developers, this is a reminder that active parameter count only describes part of the bill. A design can look efficient on paper and still spend too much time moving data or waiting for another part of the system.
What the largest benchmarks establish
Ai2 reports a configuration with 1.2 trillion total parameters and 58.36 billion active per token on 512 NVIDIA B300 GPUs. Its highest observed throughput was 858 TFLOP/s per GPU. These tests used random routing to measure systems performance rather than the quality of a trained model.
The report explicitly warns that its configurations vary in batch size, parallelism, numerical precision and recomputation policy. They do not form a controlled scaling curve. Their throughput figures are reference operating points, rather than evidence for how quickly a useful model will reach a given quality level.
A separate experiment reached 2.38 trillion total parameters using DeepEP v2. Ai2 describes this as a short capacity test rather than a complete training run. The announcement does not release a trained model of that size.
The release includes practical migration changes
The version 3.0.0 release notes include changes beyond performance. Minimum supported PyTorch rises to 2.10.0. The older model ladder API and associated training scripts have been removed, so existing orchestration built around them needs attention before upgrading.
Another fix addresses packed training data whose documents do not always end with an end of sequence token. Inferring boundaries from those tokens can merge documents and silently omit training material. The new option can use document boundary metadata instead. Default behavior stays unchanged, making the choice something operators need to inspect.
The release also changes fused attention implementation while preserving compatible parameter layouts. Its notes warn that floating point results can differ, so resuming an older run does not necessarily reproduce the same numerical path. Checkpoint compatibility alone is therefore an incomplete migration test.
Available code still needs a matching environment
The official repository provides source code under Apache 2.0 and a package distributed through PyPI. Its setup documentation lists optional dependencies for attention, grouped expert computation and lower precision training. Published container images include dependencies but leave installation of the framework itself to the user.
Ai2 says those images are regularly tested on its H100 clusters and warns that other hardware or driver combinations may need adaptation. This is usable infrastructure for technically equipped teams, with environment preparation still part of the work.
The strongest reason to study Olmo core 3 is the combination of implementation and measurement detail. Teams can inspect which costs each optimization removes, reproduce a comparable workload and check the resulting model quality before treating throughput as a savings estimate.
ByteForward’s coverage of Loopie research examines another approach to model efficiency and why measured training cost matters.
Illustration is original AI generated artwork created for ByteForward. The sculpture represents distributed expert computation and does not depict actual hardware.



