AI News

Ai2 Olmo-core 3: Open MoE Training Infrastructure

G

Gate of AI Editorial Team

Editorial & Technical Analysis

Share:
Analysis 2026-10-11 © Gate of AI

Ai2’s Olmo-core 3 is an open training infrastructure stack for scaling mixture-of-experts systems, with a design aimed at trillion-parameter-range training and more flexible experimentation across hardware.

Key Takeaways

  • Olmo-core 3 is open infrastructure for training large mixture-of-experts, or MoE, models rather than an end-user model endpoint.
  • Ai2 describes the stack as designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.
  • The system changes the earlier Olmo-core MoE approach from FSDP-based weight gathering and resharding to distributed data parallelism, or DDP.
  • Under the new design, experts remain resident on GPUs and relevant data is routed to them, avoiding repeated weight gathering.
  • The project is intended to let researchers and developers train their own MoEs, adapt the infrastructure to different hardware, and experiment with routing and parallelism.
  • Ai2’s wider infrastructure work illustrates the operating environment: thousands of NVIDIA H100, B200, and B300 GPUs, clusters ranging from 88 to 1,024 GPUs, and demand for two to three times more GPU capacity than is available.

What Is Ai2 Olmo-core 3?

Ai2 has introduced Olmo-core 3 as open, scalable training infrastructure for large mixture-of-experts systems. The project is described as the foundation for the organization’s next generation of Olmo. That next-generation Olmo is intended to use an MoE architecture and is being developed with Ai2’s largest dataset and longest context window.

The important distinction is that Olmo-core 3 is infrastructure. The verified description does not announce a released end-user model with a specified parameter count, benchmark score, training-token total, or commercial API. Instead, it presents a training system that researchers and developers can use to train their own MoEs, adapt the system to different hardware, and investigate routing, parallelism, and other parts of the training process.

This positioning places Olmo-core 3 upstream of the familiar model interface. A chat application exposes the result of a trained model. A training infrastructure stack addresses the systems work required to produce that model: distributing computation, keeping hardware occupied, moving the relevant data, and adapting the training process as model designs and accelerator environments change. The source describes Olmo-core 3 as a foundation intended to provide that flexibility.

Ai2 also frames the project as part of a broader commitment to open model development. In this approach, model weights are more useful when the infrastructure and training decisions behind them are open as well. That does not mean the context establishes a specific performance advantage, cost reduction, or production outcome. It means the project’s stated emphasis includes access to the training infrastructure and decisions behind future Olmo development.

Why MoE Training Infrastructure Matters

Mixture-of-experts models divide a model into expert components and route work among them. The verified context identifies routing and parallelism as areas that researchers can experiment with using Olmo-core 3. It also says that the stack is designed for MoE training at the trillion-parameter range while preserving computational efficiency.

That target creates a systems challenge. A large MoE is not only a model definition. It also requires a way to place experts on available accelerators and deliver the relevant data to those experts during distributed training. The architecture must therefore account for how computation is divided across devices and how the training process changes when experts are distributed across a large cluster.

Olmo-core 3 is presented as a response to that challenge rather than as a generic model-training announcement. Its stated purpose is to give researchers more flexibility as models and hardware evolve. This matters because a training stack tied too closely to one arrangement of hardware or one model design is harder to adapt. The project’s open design is intended to support experimentation with different hardware, routing approaches, parallelism strategies, and related parts of the system.

The phrase “trillion-parameter range” should be read as a description of the infrastructure’s intended scale, not as evidence that a trillion-parameter model has been released in the supplied context. The verified material does not provide a completed model’s parameter total, a training result, a benchmark table, or a claim that every organization can operate such a system. The reliable conclusion is narrower: Ai2 is building a training stack aimed at substantially larger MoE systems than its earlier work.

How Olmo-core Has Evolved

Olmo-core has changed across generations of Olmo. Ai2’s earlier sparse-model work included OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3, in contrast, used a dense architecture. In a dense architecture, nearly all of the model was active for every token, and its training stack was built around that design.

Olmo-core 3 extends the framework with a training system intended for much larger MoE models. The significance of that evolution is not simply...

Continue Reading

Log in for free to read the rest of this article and access exclusive AI tools.

Log in / Register
GateOfAI AI Guide
Online
Hello! Welcome to GateOfAI. I am your guide copilot. I can answer questions about our SaaS tools, pricing, vetted developers, and escrow safety. How can I help you today?