Dev News Daily ENDE

Ai2's Olmo-core 3 trains MoE models at 2.7 times its old throughput

The Allen Institute for AI has released Olmo-core 3, a new version of the open framework it uses to train its Olmo language models. The headline change is a redesigned system for training mixture-of-experts models, which Ai2 says is built to scale into the trillion-parameter range and is one of the core systems behind the next generation of Olmo.

Mixture-of-experts models contain many specialised components, the experts, but each token uses only a few of them. The catch, the post explains, is that the whole model still has to be stored across GPU memory and updated, and routing tokens to the right experts across a cluster adds communication costs that can erode the advantage.

The main architectural change is a move away from fully sharded data parallelism, which gathered and re-sharded the weights for every small batch, to a design based on distributed data parallelism that keeps experts resident on GPUs and sends the data to them. Around it sit expert parallelism, which spreads experts across GPUs, pipeline parallelism, which splits layers across groups of GPUs, and a distributed optimizer, which spreads optimizer state instead of copying it to every GPU. Routing metadata stays on the GPU, so the CPU can queue work without waiting for it to be copied back.

Ai2 gives two measurements. Growing the expert pool from 8 to 128 while still selecting four experts per token kept active parameters at about 3.2 billion, raised total parameters from 4.6 billion to 47 billion, and cost less than 5% of training throughput. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter model processed 52,000 tokens per second per GPU, against 19,400 with the earlier implementation, about 2.7 times more. The stack has also been benchmarked at over a trillion total parameters.

The post names NVIDIA's Megatron-Core as the established option for large MoE training; Olmo-core 3 brings an integrated alternative into Ai2's own open framework.

Ai2's Olmo-core 3 trains MoE models at 2.7 times its old throughput
Ai2's Olmo-core 3 trains MoE models at 2.7 times its old throughput — Dev News Daily

Why it matters

Training code for large sparse models has mostly lived inside a few companies or vendor stacks. An open framework with published throughput numbers gives academic and smaller labs a way to train MoE models without rebuilding the parallelism layer themselves.

Primary source
Hugging Face Blog
https://huggingface.co/blog/allenai/olmocore3