Pro Lab
One attention pattern can only track one relationship, so transformers run several heads in parallel, each with its own Q/K/V projection. Watch the same word get routed to the verb by one head and to a similar word by another, then concatenated. The 'multi-head' in every transformer.
🎬 Free preview in EN / हिं / ES — watch it, then unlock the lab to try it yourself.