Skip to main content
Instella-MoE language model, a complex neural network, achieves a 73.22 score after post-training optimization.

Editorial illustration for Instella-MoE Language Model Improves to 73.22 Score After Post-Training

AMD's Instella-MoE Hits 73.22 With Efficient MoE Design

Instella-MoE Language Model Improves to 73.22 Score After Post-Training

4 min read

AMD released Instella-MoE, a Mixture-of-Experts language model with 16 billion total parameters but only 2.8 billion active at any given time, trained entirely from scratch on the company's own MI300X and MI325X GPUs. That last detail matters: this isn't a model fine-tuned on borrowed infrastructure, it's built end to end on AMD's ROCm software stack, using the open-source Primus and Miles frameworks for both pre-training and post-training work.

The architecture leans on two specific innovations, Gated Multi-head Latent Attention and FarSkip-Collective connectivity, choices AMD says let the sparse MoE design hit competitive benchmark scores against dense and MoE models alike, including some with larger active parameter counts. Post-training pushed the model's performance to a 73.22 score, according to AMD's own reporting.

AMD is releasing the full stack of artifacts alongside the model: weights from every training stage, the data mixtures used, training configurations, and code. The company frames this as part of a broader push toward open, reproducible AI research built on AMD hardware, rather than a one-off release meant to sit behind a paywall or API.

AMD is excited to introduce Instella-MoE, a state-of-the-art fully open Mixture-of-Experts (MoE) language model with 16 billion total parameters and 2.8 billion active parameters. Trained from scratch on AMD Instinct™ MI300X and MI325X GPUs with AMD ROCm™ software stack, Instella-MoE combines a sparsely activated MoE design with architectural innovations such as Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective connectivity.

Why this matters

AMD just handed developers a fully open 16-billion-parameter model that beats Olmo3-7B-Think on average score, and it did so training entirely on its own Instinct MI300X and MI325X hardware with the ROCm stack. That's the part worth sitting with: this isn't a paper demo, it's a shipped pipeline from pretraining through SFT, DPO, and a final "Think" post-training stage, with each step's score published in a table for anyone to check. For researchers, the 71.58-to-73.22 progression is a useful reference point for how much post-training actually buys you on a sparse MoE architecture with only 2.8 billion active parameters.

For founders and infra teams, it's a signal that ROCm is now producing competitive open models, not just running someone else's CUDA-trained checkpoint. We'd still want to see independent benchmarking outside AMD's own tables before treating 73.22 as gospel, but as a fully open MoE with real architectural choices like Gated MLA, Instella-MoE is worth pulling apart yourself.

Common Questions Answered

What are the key architectural innovations in AMD's Instella-MoE language model?

Instella-MoE features two major architectural innovations: Gated Multi-head Latent Attention (Gated MLA) and FarSkip-Collective connectivity. These innovations work within a sparsely activated Mixture-of-Experts design to improve the model's efficiency and performance across various benchmarks.

How many parameters does Instella-MoE have and how many are active during inference?

Instella-MoE has 16 billion total parameters, but only 2.8 billion parameters are active at any given time during inference. This sparse activation design allows the model to maintain competitive performance while reducing computational requirements.

What hardware and software stack was used to train Instella-MoE from scratch?

Instella-MoE was trained entirely on AMD's Instinct MI300X and MI325X GPUs using the AMD ROCm software stack. The training pipeline utilized open-source Primus and Miles frameworks for both pre-training and post-training work, making it a fully end-to-end AMD infrastructure solution.

How did Instella-MoE's performance improve through the post-training process?

Instella-MoE improved from a score of 71.58 to 73.22 through a comprehensive post-training pipeline that includes SFT (Supervised Fine-Tuning), DPO (Direct Preference Optimization), and a final 'Think' post-training stage. Each step's score is published in a table for transparency and reproducibility.

How does Instella-MoE compare to other open-source language models in its class?

Instella-MoE beats Olmo3-7B-Think on average score despite being a fully open model trained entirely on AMD's own infrastructure. This achievement demonstrates that AMD's ROCm stack and Instinct GPUs can produce competitive results without relying on borrowed or external hardware resources.

LIVE20:23Cognition Buys Poke, an AI Agent for iMessage and SMS