🔒

Pro Lab

Speculative Decoding Playground

See how a small draft model makes a big model generate 2–4× faster, with provably identical output. The draft guesses the next few tokens, the target verifies them in one pass, and the emitted-token distribution locks onto the target's exactly, no matter how weak the draft. Dial draft quality and speculation length and watch the speedup.

Start 30 days free → Try 34 free labs No card needed · cancel anytime Already Pro? Log in