Back to the radar
ResearchAnnouncement

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

The essentials, the implications, and the sources behind the story.

01 / The short version

What happened

Olmo-core 3, released by the Olmo team, upgrades its open framework with a redesigned mixture-of-experts training system targeting trillion-parameter models. The publisher reports throughput gains from switching to distributed data parallelism, plus optimizations like rowwise expert parallelism and MXFP8. In publisher benchmarks on NVIDIA B300 GPUs, a 47B-parameter MoE reached 52,000 tokens/sec/GPU, and a 1.2T-parameter model hit 858 TFLOP/s/GPU. The team notes these tests used random routing and short-capacity runs, not full training quality.

See the exact references

02 / Key takeaways

What you need to know

  1. 01

    Olmo-core 3 scales MoE training to over one trillion total parameters using DDP instead of FSDP.

  2. 02

    A 47B-parameter MoE on eight NVIDIA B300 GPUs processed 52,000 tokens/sec/GPU, about 2.7x the prior implementation.

  3. 03

    MXFP8 raised end-to-end throughput ~21% versus BF16 on four B300 GPUs while cutting peak active memory from 103 GiB to 95 GiB.

Keep in perspective

What to watch for

Publisher benchmarks used random routing and short-capacity tests, so they measure system performance rather than trained model quality or sustained training.

Go to the source

Exact references

These are the original pages used for this brief. Publisher claims are not independent evaluations.

01Primary source · Hugging FaceRead the original announcementhttps://huggingface.co/blog/allenai/olmocore3

AI-generated from the linked source. It can miss context; verify consequential details in the original. How the radar works