Open-source ИИ / Брифинг по ИИ

Allen Institute for AI releases Olmo-core 3 framework for mixture-of-experts training

The Allen Institute for AI has launched Olmo-core 3, a redesigned open framework built to scale mixture-of-experts models to over one trillion parameters with minimal overhead.

Изображение, сопровождающее оригинальный отчёт на Hugging Face
Из Hugging Face. Оригинальное изображение из источника.

The Allen Institute for AI released Olmo-core 3, an open-source framework redesigned to optimize mixture-of-experts training. The system enables developers to scale models toward the trillion-parameter range while maintaining computational efficiency by selecting only a subset of specialized components for each unit of processed text.

Optimized mixture-of-experts training stack

Olmo-core 3 introduces a training system specifically architecture for mixture-of-experts (MoE) models. Unlike dense architectures that activate all parameters for every token, MoE systems use specialized components. This framework addresses the communication and coordination costs that typically rise as the number of experts increases in a distributed GPU cluster.

Internal testing showed that expanding the expert pool from 8 to 128 experts allowed total parameter capacity to grow from 4.6 billion to 47 billion. By activating only four experts per token, the training throughput decreased by less than 5% despite the significant increase in total model size. The infrastructure has been successfully benchmarked for models exceeding one trillion parameters.

Availability and technical constraints

The release includes the core code, a technical report, and an interactive demonstration. The system is designed to provide academic researchers and smaller laboratories with tools to manage the trade-offs between computation speed and data movement. This infrastructure serves as the foundation for the next generation of Olmo models.

While the framework optimizes the training process, the full model must still be stored across GPU memory and updated during training. The reported performance gains depend on balancing computation and communication across specific hardware clusters. The institute has not yet released a specific trillion-parameter model trained with this stack, only the benchmarking results for the infrastructure itself.

Первоисточник

В этом отчете кратко изложен материал указанного ниже источника. Аналитика приводится отдельно; заявления о продуктах и исследованиях атрибутируются их авторам.

Читать оригинал на Hugging Face

← Назад ко всем новостям ИИ