Mixture of Experts from Scratch — Part 2: Scaling to Billions (2017–2022)
From thousands of experts to trillion-parameter models — sparse gating, top-k routing, load balancing, the Switch Transformer, and the engineering behind scaling MoEs — all derived step by step with a concrete 4-expert example.