Pro Lab
How a trillion-parameter model runs only a fraction of itself per token. Swap one big feed-forward net for many small experts plus a gate that routes each token to its top-k. Turn k to trade compute for capacity, and watch one greedy expert force a load-balancing loss.
🎬 Free preview in EN / हिं / ES — watch it, then unlock the lab to try it yourself.