Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Blog post from Hugging Face
Ai2 has released Olmo-core 3, an open training framework redesigned to efficiently train mixture-of-experts language models at scales reaching trillions of parameters. The system replaces its earlier fully sharded approach with distributed data parallelism that keeps experts on GPUs and routes tokens to them, improving throughput in a preliminary 47-billion-parameter benchmark by about 2.7 times while allowing expert capacity to increase substantially with limited performance loss. It combines expert, pipeline, and optimizer parallelism with routing and computation optimizations such as GPU-resident routing, grouped matrix operations, and MXFP8 low-precision computation, which raised measured throughput by roughly 21% and reduced memory use in one test. Benchmarks included a 1.2-trillion-parameter configuration across 512 GPUs and a shorter capacity test reaching 2.38 trillion parameters, although these evaluated systems performance using random routing rather than trained-model quality. The accompanying report also examines practical training findings, including pitfalls in balancing expert workloads and the fact that overlapping communication with computation does not always improve speed. Olmo-core 3 will support Ai2’s next MoE-based Olmo model and is publicly available for researchers and developers to adapt, evaluate, and use for their own training experiments.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| LLM | 2 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.