🔒

Pro Lab

RLHF Reward Model Playground

How does RLHF turn 'this answer is better' into a number a model can optimize? Train a reward model from pairwise human preferences with the Bradley-Terry loss and watch it learn to score good responses above bad, recovering the hidden notion of quality from comparisons alone.

Start 30 days free → Try 34 free labs No card needed · cancel anytime Already Pro? Log in