IA3 MIN

Olmo-core 3 helps larger AI models train faster by changing how the work moves

Ai2 has detailed an open training system that cuts repeated data transfers. One test expanded a model to ten times its size while losing less than 5% of its training speed.

One token selects four expert modules, highlighted in green within a grid of 128. Other model parts stay active, and a batch of tokens can use all 128 experts. This is a conceptual example, not a fraction of all model parameters.
Image: INSERT FUTURE · gráfico original / original graphic. Datos y mecanismo / data and mechanism: Ai2.

Ai2 has detailed Olmo-core 3, an upgrade to its open software for training AI models. It tackles a problem that can make larger models expensive to train even when they use only a small portion of their components at a time: moving data around the chips doing the work. Researchers can inspect, adapt and reuse the implementation.

This is infrastructure for building models. Training adjusts the numerical values that determine how a model processes information, a distinction covered in our guide to training and fine-tuning. Olmo-core coordinates those calculations across GPUs. The release does not provide a new consumer chatbot.

01

A bigger expert pool without using every expert

A mixture-of-experts model splits some of its computation among separate modules. A routing component chooses which modules handle each token, a small unit of text. Adding more modules increases the model’s total capacity without requiring every token to use them all. Across a batch of text, however, all the experts may do work.

Training still has to store and update the whole model. The earlier system repeatedly gathered and redistributed model values for small batches of training data. Olmo-core 3 instead keeps expert partitions on their GPUs while accumulating work before an update, and routes the relevant data to them. It reduces repeated transfers without eliminating the coordination and updates that training requires.

In the old system compared here, weight shards are gathered every microbatch. In the new design, tokens travel to experts whose weights stay in their assigned GPUs between updates. Both still coordinate model updates; the three GPUs are schematic.
Image: INSERT FUTURE · gráfico original / original graphic. Datos y mecanismo / data and mechanism: Ai2.
02

What Ai2 actually measured

One experiment in Ai2’s technical report used eight NVIDIA B300 GPUs and expanded the available expert pool from 8 to 128. Each token still selected four experts. Total model size grew from 4.6 billion to 47 billion parameters, the numerical values it learns, while approximately 3.2 billion remained active per token. The new system processed 54,500 and 52,000 tokens per second per GPU at those two sizes: a 4.6% slowdown for roughly ten times the total capacity.

At 47 billion parameters, the earlier implementation managed 19,400 tokens per second per GPU. The new stack’s 52,000 is about 2.7 times that throughput. Several components changed together, so this compares the two Olmo-core implementations rather than isolating one optimization or establishing a lead over every rival framework.

The test used routing that approximated balanced workloads. Real training data can send more work to some experts than others and reduce throughput. These measurements also say nothing by themselves about the quality of a trained model’s answers or the total bill for a completed training run.

Preliminary whole-system comparison with about 3.2 billion active parameters. As total parameters rise from 4.6 to 47 billion, old-stack throughput falls from 50,500 to 19,400 tokens per second per GPU; the new stack goes from 54,500 to 52,000. Eight B300 GPUs, eight layers and balanced random selection of four experts per token.
Image: INSERT FUTURE · gráfico original / original graphic. Datos y mecanismo / data and mechanism: Ai2.
03

Available code, with a future model still to come

The official repository makes the code available under Apache 2.0. Its changelog dates version 3.0.0 to September 30, one day before Ai2’s detailed announcement. Ai2 plans to use the system for its next generation of Olmo. Researchers can study and adapt the training machinery now, but the current results do not evaluate that future model’s capabilities.

00

The conversation starts here

Sign in with a supporter account to comment. Sign in

Nobody has commented yet. Want to go first?

KEEP READING

You may also like

FRONT PAGE