Mixture of Experts Explained: The Math That Makes AI Cheaper
Mixture-of-Experts (MoE) is the architecture behind why frontier models keep getting bigger on paper while getting cheaper to run in practice. This piece breaks down total parameters versus active parameters per token, using DeepSeek-V3, Llama 4, Qwen3, and Kimi K2 as verified examples. The trend line matters more than any single spec: active-parameter share has been shrinking with every major release since 2023, decoupling capability from compute cost. For PMs, a vendor's headline parameter count is now closer to a marketing figure than a cost or performance signal, and knowing which number to ask for is a negotiating advantage.