🔒

Pro Lab

The KV-Cache

Why LLM generation is O(n) not O(n²). Watch autoregressive decoding re-encode the whole prefix every step (the wasteful triangle) versus caching each token's keys/values once (the diagonal), and the memory-for-speed trade that caps context length.

Start 30 days free → Try 34 free labs No card needed · cancel anytime Already Pro? Log in