All learning paths Learn See Try Make What a reward model is What this lesson is about See what is built from marked pairs: a scorer, trained on your order rather than on a right answer.
Your batch of eight hundred marked pairs goes to a lab on Monday. A colleague asks what the lab builds from it.
A The lab keeps the chosen reply of each pair as the one correct reply for that request. B A reward model is a scorer trained on your pairs, not on the right answer. C Both replies of every pair are fed into the training run, so the rejected one teaches too.
Check my recall Why this matters A published summarisation run shows the shape. Marked pairs train a scorer, and the model is then improved against that scorer. So your pair is not an answer the model copies. It is one line in the scorer's own teaching. The model is later pushed to please that scorer.
Say what a reward model is, what it is trained on, and what one marked pair becomes.
A published summarisation run shows the shape. Marked pairs train a scorer, and the model is then improved against that scorer. So your pair is not an answer the model copies. It is one line in the scorer's own teaching. The model is later pushed to please that scorer.
A reward model is a scorer trained on your pairs, not on the right answer. One pair teaches order only: this reply outranks that one, for that request. The model is later pushed to raise that score, so a weak scorer teaches weak habits. Task Eight hundred marked pairs of health replies go for training. Say what one pair becomes.
Weak approach Says each chosen reply becomes a worked example for the model. Reads the eight hundred as eight hundred ideal answers the model will copy.
Stronger approach Says each pair becomes one comparison for a scorer. The scorer learns that one reply outranks the other. The model is then pushed to raise that score.
Why the stronger approach works The weak reading turns the chosen reply into an ideal answer to copy. The pair carries order, not wording.
Try a changed situation A shop-review team marks Marathi pairs. Its scorer now puts the longer reply on top almost every time.
Make something yourself Take one pair you marked this week. Write what a scorer learns from it, and one thing it cannot learn from it.
Lesson 6 of 9 on How post-training works next Direct preference training Part 2 ends in a work sample you can send
Read the primary or official source: Stiennon et al., Learning to summarize from human feedback (arXiv)