01 / The short version
What happened
Olmo-core 3, released by the Olmo team, upgrades its open framework with a redesigned mixture-of-experts training system targeting trillion-parameter models. The publisher reports throughput gains from switching to distributed data parallelism, plus optimizations like rowwise expert parallelism and MXFP8. In publisher benchmarks on NVIDIA B300 GPUs, a 47B-parameter MoE reached 52,000 tokens/sec/GPU, and a 1.2T-parameter model hit 858 TFLOP/s/GPU. The team notes these tests used random routing and short-capacity runs, not full training quality.
See the exact references02 / Key takeaways
What you need to know
- 01
Olmo-core 3 scales MoE training to over one trillion total parameters using DDP instead of FSDP.
- 02
A 47B-parameter MoE on eight NVIDIA B300 GPUs processed 52,000 tokens/sec/GPU, about 2.7x the prior implementation.
- 03
MXFP8 raised end-to-end throughput ~21% versus BF16 on four B300 GPUs while cutting peak active memory from 103 GiB to 95 GiB.
Keep in perspective
What to watch for
Publisher benchmarks used random routing and short-capacity tests, so they measure system performance rather than trained model quality or sustained training.
Go to the source
Exact references
These are the original pages used for this brief. Publisher claims are not independent evaluations.
01Primary source · Hugging FaceRead the original announcementhttps://huggingface.co/blog/allenai/olmocore3AI-generated from the linked source. It can miss context; verify consequential details in the original. How the radar works