Ai2 has released Olmo-core 3, an open framework for developing large language models with a redesigned mixture-of-experts (MoE) training system. The framework is intended to scale MoE training toward the trillion-parameter range while keeping computation efficient, and will serve as core infrastructure for the next generation of Olmo. In one benchmark, expanding the expert pool from 8 to 128 while routing each token to four experts increased total capacity from 4.6 billion to 47 billion parameters, with active parameters per token held near 3.2 billion and throughput declining by less than 5%. A 47-billion-parameter MoE reached 52,000 tokens per second per GPU on eight NVIDIA B300 GPUs, compared with 19,400 for the earlier FSDP-based implementation, or about 2.7 times the throughput. Olmo-core 3 replaces repeated weight gathering and resharding with a distributed data-parallel design that keeps experts on GPUs and routes data to them. It combines expert, pipeline and distributed-optimizer parallelism with rowwise expert parallelism, GPU-resident routing and grouped GEMM to reduce memory, communication and computation overhead. In a four-GPU benchmark, MXFP8 raised end-to-end throughput by about 21% over BF16 and reduced peak active memory from 103 GiB to 95 GiB. The system was benchmarked with a 1.2-trillion-parameter model across 512 GPUs, reaching 858 TFLOP/s per GPU, while a separate short-capacity test reached 2.38 trillion total parameters using DeepEP v2. These large-scale tests used random routing or limited capacity tests, so they demonstrate infrastructure scale rather than trained-model quality or sustained full-training performance. The accompanying report also documents cautions about routing metrics, expert learning rates, input-dependent GPU timing and communication-computation overlap. The code, technical report and interactive walkthrough are open for researchers and developers to adapt to different hardware and experiment with MoE training.
