🔒

Pro Lab

Mixture of Experts

How a trillion-parameter model runs only a fraction of itself per token. Swap one big feed-forward net for many small experts plus a gate that routes each token to its top-k. Turn k to trade compute for capacity, and watch one greedy expert force a load-balancing loss.

🎬 Free preview in EN / हिं / ES — watch it, then unlock the lab to try it yourself.

Start 30 days free → Try 34 free labs No card needed · cancel anytime Already Pro? Log in