Skip to lesson
  1. Learn
  2. See
  3. Try
  4. Make

What a reward model is

What this lesson is about

See what is built from marked pairs: a scorer, trained on your order rather than on a right answer.

Your batch of eight hundred marked pairs goes to a lab on Monday. A colleague asks what the lab builds from it.

Why this matters

A published summarisation run shows the shape. Marked pairs train a scorer, and the model is then improved against that scorer. So your pair is not an answer the model copies. It is one line in the scorer's own teaching. The model is later pushed to please that scorer.

Lesson 6 of 9 on How post-training worksnextDirect preference trainingPart 2 ends in a work sample you can send

Read the primary or official source: Stiennon et al., Learning to summarize from human feedback (arXiv)