All learning paths Learn See Try Make Direct preference training What this lesson is about Compare the two routes: train a scorer and use RL, or move the model straight from the pair.
A firm tells your circle it will no longer build a scorer for this work. Nobody in the room knows what that means.
A The scorer will be built at another desk and handed to the firm with the batch. B Direct preference training moves the model straight from a chosen and a rejected reply. C With nothing to score the pairs, your marks are only advice now, so the batch can wait.
Check my recall Why this matters The older route trains a scorer from your pairs, then improves the model against it with RL. Direct preference training, published in 2023, drops the scorer and moves the model straight from the chosen and rejected reply. Your marking does not change. One habit now costs more: a pair marked as a tie gives this route nothing at all.
Say how direct preference training differs from the older two-step route, and what it asks of your batch.
The older route trains a scorer from your pairs, then improves the model against it with RL. Direct preference training, published in 2023, drops the scorer and moves the model straight from the chosen and rejected reply. Your marking does not change. One habit now costs more: a pair marked as a tie gives this route nothing at all.
Direct preference training moves the model straight from a chosen and a rejected reply. The older route needs two steps: teach a scorer, then improve the model against it with RL. The number to check in your own batch is how many pairs you marked as a tie. Task A buyer asks your circle to re-check three hundred old pairs for a direct preference run.
Weak approach Says the old marks will do, because the replies never changed and the better one is already ticked. Sends all three hundred on without opening the sheet.
Stronger approach Counts the ties first, and finds forty-one of three hundred. Sends those back to be decided or dropped, and writes the reason on the sheet.
Why the stronger approach works Both re-read the same batch. Only the strong one counts the marks this kind of run cannot use.
Try a changed situation Your three hundred pairs carry forty-one ties. The buyer wants a direct preference run this week.
Make something yourself Take a batch of pairs you marked. Report how many are ties, and say what you would do with them before a direct preference run.
Lesson 7 of 9 on How post-training works next Where your labels go Part 2 ends in a work sample you can send
Read the primary or official source: Rafailov et al., Direct Preference Optimization (arXiv)