GitHub TrendingGitHub Trending 5 min read

Ai2 Opens Olmo-core 3 for Trillion-Scale MoE Training

Allen Institute for AI released Olmo-core 3 on Thursday, 1 October 2026, a framework SiliconANGLE says can train mixture-of-experts models at trillion-parameter scale.

PC

PromptCrates Editorial

Staff Writer

0 0
Ai2 Opens Olmo-core 3 for Trillion-Scale MoE Training

Allen Institute for AI released Olmo-core 3 on Thursday, 1 October 2026, a development framework that SiliconANGLE reports can train mixture-of-experts large language models at trillion-parameter scale while keeping throughput high enough for research labs that lack hyperscaler budgets. In Ai2’s benchmarks, Olmo-core 3 processed about 52,000 tokens per second on Nvidia B3000 GPUs for a 47 billion-parameter model, roughly 2.7 times the approximately 19,400 tokens per second Ai2 cited for Nvidia’s Megatron-core.

What MoE training usually costs

Mixture-of-experts models split computation across specialist subnetworks for each token instead of activating every parameter the way dense models do. That design can pack far more total parameters into a system while computing with only a few experts at a time, which lowers inference and training FLOPs relative to a dense model of the same total size. The catch, SiliconANGLE notes, is that the full expert pool still has to live across GPU memory, and coordinating experts during training creates networking and optimizer overhead that often erases the paper savings.

Ai2 said Olmo-core 3 was built to bridge dense and MoE workflows by growing the expert pool from eight to 128 while still selecting only four experts per token, then scaling the same infrastructure past one trillion parameters. The white paper summarized by SiliconANGLE describes expert parallelism that spreads experts across multiple GPUs so each card stores only part of the pool, splitting model layers across GPU groups to cut per-card memory, and a distributed optimizer that shards optimizer state instead of replicating full copies everywhere. Ai2 also supports MXFP8, a lower-bit number format that can reduce computation and the volume of data moved between GPUs.

Why the throughput number matters for open research

A 2.7 times throughput jump against Megatron-core is not a leaderboard flex for a finished chat model; it is a training-time cost claim. For university groups and independent labs, wall-clock tokens per second decide whether a trillion-parameter MoE experiment finishes before a grant cycle ends. Ai2’s stated vision, per SiliconANGLE, is to put larger MoE training within reach of researchers who cannot rent enterprise-scale clusters, and to let them experiment with routing, parallelism, and hardware adaptations rather than treating those knobs as hyperscaler secrets.

That open-research framing fits Ai2’s broader Olmo line of transparent models and tooling. PromptCrates has tracked neighboring open releases such as AWS’s Strands Decider 2B for fast agent choices and Hindsight’s agent-memory stack on GitHub, where the news value is not only capability but the fact that working code landed in public repositories. Olmo-core 3 continues that pattern: SiliconANGLE says the project and related systems are available on GitHub for developers and the open-source community now.

Limits buyers and researchers should still test

Benchmark claims always need independent reproduction. Ai2’s 52,000 tokens-per-second figure is tied to a specific 47 billion-parameter configuration on B3000 hardware; labs on older GPUs, different interconnects, or denser routing schedules may see smaller gains. Expert parallelism and distributed optimizers also introduce failure modes—stragglers, routing collapse, numerical instability under MXFP8—that only show up when someone else trains for more than a smoke test. Researchers should treat the white paper’s architecture diagram as a starting checklist: confirm expert sharding, layer grouping, optimizer state placement, and number formats on their own clusters before budgeting a trillion-parameter run.

For companies already standardized on Megatron-core or other stacks, Olmo-core 3 is a competitive pressure point more than an overnight swap. Training frameworks stick because of operational muscle memory, checkpoint formats, and monitoring. Ai2’s contribution is to show that MoE efficiency is still an open contest, not a closed Nvidia-only story, and that a nonprofit research institute can publish the knobs. If independent groups reproduce the throughput claim, expect more MoE ablations from labs that previously stayed dense because the tooling tax felt too high.

Documentation quality will decide adoption as much as the 2.7 times throughput claim. Researchers need reproducible recipes: exact parallel degrees, batch sizes, routing temperatures, and failure modes when an expert shard drops mid-run. If Olmo-core 3 ships with those runbooks beside the GitHub tree, university clusters can copy a known-good configuration instead of reverse-engineering Ai2’s internal cluster. If the release is mostly code without operational narrative, the framework risks becoming a citation rather than a daily trainer.

There is also a policy angle. Open MoE tooling lowers the barrier for labs outside the United States and outside cloud credits programs, which can diversify who trains large sparse models and who publishes ablations. That diversity is healthy for science and complicated for export-control conversations that already orbit frontier training runs. Ai2’s nonprofit posture makes Olmo-core 3 easier for many institutions to adopt than a vendor SDK with opaque telemetry, but IT security teams will still ask what phone-home endpoints exist and how checkpoints are hashed before they open firewall rules to the training fabric.

Primary reporting for this article: Kyt Dotson’s SiliconANGLE coverage of Ai2’s 1 October 2026 Olmo-core 3 release, including the expert-pool scaling claim, the 52,000 versus 19,400 tokens-per-second comparison, and the expert-parallelism, layer-split, distributed-optimizer, and MXFP8 details drawn from Ai2’s white paper.

github-trendingAi2OlmoMoEopen-source

Related articles