AI RundownDaily
Topic

#inference cost

3 articles — updated daily

Model Distillation Explained: Why Every AI Lab Ships It

Model Distillation Explained: Why Every AI Lab Ships It

Model distillation is the process of training a smaller "student" model to mimic a larger "teacher" model's outputs, and by mid-2026 every major AI lab ships a distilled sibling alongside its flagship. The pattern spans GPT-5.4 Mini, Gemini 3.1 Flash/Flash-Lite, Claude Haiku, and open releases like DeepSeek-R1-Distill. The bigger signal is that distillation has become the default architecture for shipping AI at production scale, not a discount option. For PMs, the real work is knowing when a distilled model quietly costs you accuracy versus when it's the obviously correct call.

Mixture of Experts Explained: The Math That Makes AI Cheaper

Mixture of Experts Explained: The Math That Makes AI Cheaper

Mixture-of-Experts (MoE) is the architecture behind why frontier models keep getting bigger on paper while getting cheaper to run in practice. This piece breaks down total parameters versus active parameters per token, using DeepSeek-V3, Llama 4, Qwen3, and Kimi K2 as verified examples. The trend line matters more than any single spec: active-parameter share has been shrinking with every major release since 2023, decoupling capability from compute cost. For PMs, a vendor's headline parameter count is now closer to a marketing figure than a cost or performance signal, and knowing which number to ask for is a negotiating advantage.

Together AI's $800M Round Signals Open-Source AI's Rise

Together AI's $800M Round Signals Open-Source AI's Rise

Together AI raised $800 million at an $8.3 billion valuation, according to Reuters, one of the largest bets yet on open-source AI infrastructure. The round signals that investors see real, durable demand in helping companies run open-weight models outside the big proprietary labs. It's a sign that vendor lock-in has become a boardroom-level risk, not just an engineering preference. For PMs, it means the case for open-source model infrastructure, on cost and control, just got a lot harder to ignore.